Last December, a strange phone went on sale in China. It could open your apps, find the right buttons, tap them for you. Here was an AI assistant that could actually control your phone. Only 30,000 units existed. They sold out in 48 hours, and resellers soon asked around $1,110 for a phone that retailed at $490.
Three weeks later it was a brick. Not because it broke. Because WeChat started flagging its logins as abnormal, Alipay threw up verification walls, and banking apps simply refused to serve it. Owners of the most hyped phone of the season found that their apps had quietly gone on strike.
The phone was Nubia’s M153, the first device with ByteDance’s Doubao assistant built in. The lesson from its rise and collapse is the most useful thing that has happened in consumer AI this year. And this week, the second attempt shipped with that lesson written into a contract.
Three ways to give an AI your phone
There are three known ways to let an assistant control your apps. The first is to let it see the screen and tap like a human does, which engineers call a GUI agent. It is universal, because every app gets controlled the same way. It is also fragile and, as the M153 demonstrated, trust-free. From the app’s side, an assistant tapping buttons looks exactly like a bot, and bots get blocked.
The second way is official interfaces. Apple’s App Intents and the equivalent machinery on Android let an app expose specific actions it is willing to have performed on its behalf. This is stable, auditable, safe. It is also nearly empty, because app makers have little reason to build hooks for an assistant that might route customers around their screens. Every AI assistant sold in America today lives inside this bottleneck.
The third way is what Doubao just tried. The consumer version launched September 14 on the new Nubia NaviX Ultra, and it runs the GUI approach as a labeled beta. Underneath sits something new: a protocol called SAEP, which gives every app a formal way to say no.
The deal is the product
SAEP is short for Screen Automation Execution Protocol, and its rules matter more than any benchmark. Apps can opt out entirely, or allow automation but forbid specific actions like posting, deleting, or claiming rewards. There is a 30-day public comment period. Anything that has not said yes stays off limits by default.
So the assistant’s ceiling is now a matter of negotiation. WeChat, for instance, currently allows voice messages and calls through the assistant but not payments or transfers. That is the deal in miniature. The AI can do what the app concedes. Nothing more.
A polite assistant with a low ceiling has replaced a capable one that got everyone’s account frozen. My judgment is that this trade is correct, and the first generation proved it by failure. Technical capability was never the bottleneck. The M153 could genuinely complete tasks before the blocks landed. What it could never get was permission, and permission is not a feature you can ship.
But there is a real cost, and it should be named. Every capability the protocol returns to the assistant was taken by an app’s decision, which means the app makers now collectively set the pace of phone AI. If the most important apps in China decide to say no, Doubao’s second act ends as a very good voice assistant with a fingerprint button. The GUI agent dream waits another cycle.
Why did this collision happen in China first, and not in California? Partly by accident, partly by structure. Chinese super apps concentrate payments, chat, shopping, and government services inside a handful of companies, so a rogue agent threatens everything at once and gets crushed instantly.
American app usage is a long tail of hundreds of smaller apps, where the same threat is diluted across a thousand small conflicts. The protocol arrived where the pain was loudest. That is structure, not virtue.
It also will not play out the same way in the United States. American app makers are more litigious, app store policies are stricter about automation, and the platforms themselves, Apple and Google, prefer to own the assistant layer through official interfaces.
Expect the American path to stay on route two for years: safe, slow, limited to whatever apps bother to expose. The Chinese path, messy and protocol-based, will cover more ground faster precisely because it tolerates the mess.
Two cold showers before the hype sets in, though. The company’s own figure for end-to-end task reliability is 80%, which is an honest-sounding number with a flexible definition attached. And the new hardware button that wakes the assistant, secured with a fingerprint reader, is a reminder of who holds the real switch.
It is held by the phone maker. Not by you. And not by the apps.
The local memory side deserves its own paragraph. The assistant indexes your photos, messages, and notes on the device, and you can inspect, edit, or wipe what it keeps. That is a bigger privacy shift than any cloud assistant has offered, and it is also the feature most likely to make a regulator nervous.
On-device memory of everything you have ever said to your phone is either the most private design possible or the juiciest breach target. Both things are true. The difference is one lawsuit away from mattering.
There is also the small matter of what happens when both ends are AI. Early testers found the assistant can answer phone calls on your behalf, and at least once it ended up in a conversation with someone else’s automated receptionist.
Two voice agents politely negotiating a restaurant reservation is funny once. As a preview of the next decade of phone traffic, it is something else.
Doubao itself is not a niche experiment. The assistant sits on a base app with around 168 million monthly users, which is more people than live in Russia. ByteDance can afford to lose hardware money for years to win the entry point, the same patience that made Chinese token prices a standing talking point in Silicon Valley.
I wrote before about how Chinese AI apps reached massive scale in silence. This is the next step of that story arriving on a phone you cannot buy, running apps you have never used, solving a problem every platform on earth will face.
Because the problem is universal. Every phone platform now ships an AI assistant, and every one of them is stuck at the same door.
Apps do not trust agents. Agents cannot function without trust. Users will not grant trust to something that gets their accounts banned. China’s answer was to write the refusal rules down and make them public, on a 30-day clock, before anything ran.
One more wrinkle worth tracking. The 30-day comment period turns app developers into lobbyists, and the loudest ones will shape the default rules for everyone else. Regulation by negotiation is quieter than regulation by court. It can also freeze out the small apps that never get a seat at the table.
My prediction: some American platform ships an SAEP-style clause in its developer policy within two years, and nobody will remember where the idea came from. Whoever copies that move first will have learned the actual lesson of the brick that sold out in 48 hours. Trust is not a model capability. It is a contract, and contracts have to be written down.