Talking to My Phone Like It's a Person Now

Talking to My Phone Like It's a Person Now

Tech News ai chatgpt openai voice assistants

So the update finally showed up in my ChatGPT app yesterday, that little "new" badge next to the voice icon, and I've been poking at it ever since instead of doing literally anything productive. OpenAI started rolling out Advanced Voice Mode to Plus and Team subscribers in the US this week, and after two days of it sitting on my phone doing nothing, mine switched on Tuesday.

If you haven't seen it yet: it's not the old voice mode, the one that transcribes what you say, sends it off as text, and reads back a reply in a slightly stilted TTS voice with a full second of dead air in between. This is actually fast. You talk, it responds almost immediately, and you can interrupt it mid-sentence and it'll just stop and listen, the way an actual person would when you cut them off. There are five voices to pick from now — I ended up on the one called Cove, mostly because Breeze sounded a little too much like a hotel concierge for my taste. Sky is gone, for the reasons anyone who was paying attention back in May already knows about, and I won't rehash that here.

I made coffee this morning while asking it to walk me through recalculating the ratio because I'd run out of my usual beans and grabbed a bag that was way stronger than what I'm used to. It kept up, adjusted when I told it I only had a 6-cup pour-over and not the 8-cup I'd mentioned a minute earlier, and didn't miss a beat when my downstairs neighbor's dog started losing its mind at the mailman halfway through and I had to shout over it. That's the part that actually impressed me, not the voice quality, which is good but not magic, but the fact that it handled being talked over and talked around like a conversation instead of a transcript.

Here's my actual complaint though, and it's not a small one: it still can't see anything. No camera, no screen share, none of it. So despite feeling like a real-time conversation, it's still just a phone call with something that reads very well and has a nice cadence. Half the reason I was excited for this originally, going back to the demo back in the spring, was the bit where it looked at a math problem on paper and walked through it, or glanced at a room and described it. None of that is live yet. So what I actually have right now is a conversational layer bolted onto the same text-only brain, and the gap between what it felt like it should do and what it does do is a little deflating once the novelty wears off, which took about forty minutes for me.

It's also US-only for now, no UK, no EU, no Switzerland, not even Iceland or Norway got it in this wave, apparently for regulatory reasons OpenAI hasn't fully spelled out. I've got a few readers who've emailed over the years from Germany and one from Dublin, and if that's you, sorry, you're waiting a bit longer, and I genuinely don't know how much longer.

Funny thing is I went digging through my own archives after this to see if I'd written anything about Siri back when it first showed up, because this blog's old enough that I actually did, there's a post from October 2011 where I called it "genuinely useful about 30% of the time and a party trick the rest," which, thirteen years later, honestly still describes a lot of voice assistants pretty well, including this one on a bad day. The difference now is the 30% keeps creeping up, and the party-trick 70% is a lot more entertaining while it fails.

Anyway. If you're a Plus subscriber in the US and you don't see it yet, it's apparently a staggered rollout, so give it a few days before you start assuming your account got skipped. And if you do get it, ask it to do something conversational and a little weird, not a task, that's where it actually shows what's different. Asking it to summarize an email is a waste of the good version.