Gemini is neat, but it SSUUUUUUUUCCCKKS at being a voice assistant.
It will frequently produce a fucking essay in response to being asked to turn the lights on or off.
I just gave it a go again, 30% of the time it pretends to do what you ask without actually doing it. 20% of the time it flat-out refuses while never failing to mention that I really should enable gemini activity tracking. 40% of the time is claims this functionality does not exist, even though it was given the ability to “ask assistant” to do things only it can do (like identify music by listening to it). 10% of the time it will successfully hand off to assistant, which will then instantly do as asked.
The problem is Google is treating all those functions (“tool calling”) as just any other response, they didn’t even bother training the model to recognize and process the requests properly using the tools. They did the Google thing of just putting the function in and calling a day and forgetting about it.
So without proper reinforcement learning it responds with the same casual randomized answer-shaped sequence of words as ever, you’re just lucky if the part of it that knows to look up the tool call instructions responds, and luckier if it does it right AND submit the tool call formatted right. If it forgets to look up it’s tool call instructions, it might randomly tell you it can’t do it because you got an answer from a part of the network trained on data from when there weren’t models that could.
OMFG.
Gemini is neat, but it SSUUUUUUUUCCCKKS at being a voice assistant.
It will frequently produce a fucking essay in response to being asked to turn the lights on or off.
I just gave it a go again, 30% of the time it pretends to do what you ask without actually doing it. 20% of the time it flat-out refuses while never failing to mention that I really should enable gemini activity tracking. 40% of the time is claims this functionality does not exist, even though it was given the ability to “ask assistant” to do things only it can do (like identify music by listening to it). 10% of the time it will successfully hand off to assistant, which will then instantly do as asked.
The problem is Google is treating all those functions (“tool calling”) as just any other response, they didn’t even bother training the model to recognize and process the requests properly using the tools. They did the Google thing of just putting the function in and calling a day and forgetting about it.
So without proper reinforcement learning it responds with the same casual randomized answer-shaped sequence of words as ever, you’re just lucky if the part of it that knows to look up the tool call instructions responds, and luckier if it does it right AND submit the tool call formatted right. If it forgets to look up it’s tool call instructions, it might randomly tell you it can’t do it because you got an answer from a part of the network trained on data from when there weren’t models that could.
Fixed that for you
Gemini
is heat, but itSSUUUUUUUUCCCKKS