So Google finally shipped Gemini this week. Announced Tuesday, and by Wednesday morning my Twitter (or whatever Im supposed to call it now, still cant bring myself to type the other name) was wall to wall with that "hands-on" demo video. Youve probably seen it too - the one where someone talks to the model, draws a duck, moves cups around, and Gemini just... follows along, cracking jokes, catching the rubber duck squeak game. Its genuinely impressive to watch. My first reaction watching it at like 11pm was something close to "okay, this is the first AI demo thats actually made my stomach drop a little."
Then I went and actually read the blog post underneath the video instead of just watching it twice more, and my stomach unclenched a bit.
Buried down in Googles own post is a line that basically says the video was made by prompting Gemini with still image frames pulled from footage, then having it respond to those frames. Not live. Not continuous. Not the model watching video and reacting in real time the way the whole thing is edited to make you believe. Its a demo built after the fact out of a bunch of individual screenshots, then cut together with narration and sound design so it feels like one continuous conversation.
I dont think thats necessarily a lie, exactly. Every product demo since the dawn of product demos has been staged to some degree. Steve Jobs had a second, working iPhone hidden under the podium in 2007 in case the wifi cut out mid-demo. But theres a difference between smoothing over a live demo and building an entire minute-and-a-half video out of static frames and editing it so it plays back like it happened live. Ones stagecraft. The others closer to what a movie trailer does, showing you the version of the product that doesnt quite exist yet.
And the actual announcement underneath all this is a big deal regardless of the video. Three model sizes, Ultra, Pro, and Nano, with Ultra apparently beating GPT-4 on something like 30 of 32 benchmarks Google ran, including crossing 90% on the MMLU test thats supposed to represent expert-level knowledge across a pile of subjects at once. Bard is already running on Gemini Pro as of this week here in the US, in English, and you can feel its a step up if you poke at it for ten minutes. Ultra itself isnt out yet, they say early next year, after more safety testing.
But I keep coming back to the video thing because its such a specific, avoidable kind of mistake. Nobody needed the demo to be fake-continuous. The real capabilities, described plainly in a paragraph, are already worth writing about. Instead we got a viral clip that a decent chunk of tech coverage treated as basically real-time multimodal reasoning for about a day, before people actually read past the video and started walking it back. Thats the gap that bugs me, not the underlying model.
Ive been doing this blog since November 2011, which means Ive now sat through roughly four separate this-changes-everything demo videos and watched the same week play out every time. The original Siri ads. The first Google Duplex phone call. That Boston Dynamics parkour robot everyone assumed had to be faked (it wasnt, for the record, that one actually held up under scrutiny). The pattern never really changes: incredible video, quiet caveat eleven paragraphs down, three days of people arguing in the replies about whether the caveat matters. It always matters a little. Its never quite as damning as the skeptics want it to be, and never quite as clean as the hype wants either.
Anyway. Ill actually use the thing once its rolled out more broadly and see how Bard-with-Gemini feels for the boring stuff I do sixty times a day, summarizing a PDF, drafting a reply, sanity-checking a paragraph, not duck drawings. Thats the test that matters to me, not a trailer edited by someone in marketing. If its genuinely better than what Bard was doing last month, which honestly was not setting a high bar, Ill come back here and say so.