Alibaba is supplying the agent layer rather than building the phone
Qwen Intelligence was introduced on September 22 at Alibaba’s Apsara Conference in Hangzhou.
Alibaba describes it as a full-stack smartphone agent solution for device manufacturers, built around Qwen models and the infrastructure required to plan and execute mobile tasks.
The company explicitly says the strategy is collaborative rather than a move into manufacturing its own smartphones.
HONOR is the first major hardware partner attached to a commercial launch.
Magic9 is scheduled to be the first retail implementation
HONOR will unveil the Magic9 series in Beijing on September 28, with MagicOS 11 and Qwen Intelligence integrated into the new software platform.
HONOR’s Robot Phone project will support the same technology, but Magic9 is the first conventional consumer smartphone identified for the rollout.
The companies had already announced cooperation earlier in 2026. This week’s announcement turns that collaboration into a named platform tied to an actual shipping product.
Qwen splits planning, operation and creation into separate systems
The architecture is built around three specialized parts rather than one model being expected to handle every job.
Mobile Planner Agent interprets the user’s objective, decomposes it into smaller steps, selects tools and changes the plan when necessary.
A separate operation system is responsible for interacting with smartphone interfaces, including launching apps, selecting controls and filling fields.
A third component focuses on content creation and multimodal generation.
That separation is sensible because understanding a request and reliably pressing the correct control inside an app are fundamentally different problems.
Cross-app execution is where agents become difficult
Traditional assistants are already competent at launching one application or returning information.
A request involving travel planning might require a browser, booking service, maps, calendar, payment workflow and messaging app.
Every service introduces its own interface, confirmations and possible errors.
The agent has to preserve the original goal while reacting to whatever each step produces.
HONOR says YOYO can now handle more than 100 operations
HONOR and Alibaba say the system can execute long tasks extending beyond 100 individual actions.
The companies describe roughly 700 internal tools, compatibility with more than 500 skills and over 40 types of conditional triggers in the broader YOYO system.
That moves significantly beyond a predefined voice shortcut performing two or three known commands.
It also creates a larger failure surface. A hundred actions are a hundred opportunities to misunderstand a screen, lose context or perform something the user did not intend.
A 91.8% success number looks different when failure has consequences
The companies report 91.8% overall task accuracy, 87% accuracy on certain complex tasks and an end-to-end completion rate above 90%.
Alibaba also cites 3.6-second GUI operation performance in its testing.
Those numbers come from Alibaba and HONOR. There is not yet an independent public benchmark describing the exact app set, task distribution and exception handling behind them.
A 91.8% result can be impressive for an experimental agent. It is less reassuring if the remaining failures can purchase the wrong ticket, message the wrong contact or alter the wrong appointment.
Action changes the cost of an AI mistake
A chatbot can produce a bad answer and leave a human to decide whether to trust it.
An agent with permission to act can convert a bad answer into a real-world outcome.
That makes confirmations, previews and reliable undo mechanisms essential rather than optional interface polish.
HONOR’s design decisions around those safeguards may ultimately matter more than which Qwen model sits behind them.
Phones contain exactly the context an agent wants
Alibaba argues that smartphones are unusually suitable for personal agents because they already contain messages, photos, calendars, apps, location data and payment tools.
That context is what can make a phone agent more useful than a model isolated inside a chat window.
It also raises the privacy stakes immediately.
The more context an agent can inspect and the more applications it can control, the more important it becomes to know which information leaves the device and how cloud processing is handled.
Qwen Intelligence is not being presented as fully offline
Alibaba describes a platform spanning models, agents and cloud infrastructure rather than a completely local assistant.
Some execution can clearly happen through MagicOS and the phone’s own hardware, while heavier planning or generative tasks may call remote services.
The balance between local and cloud processing needs much more detailed documentation when Magic9 launches.
It will also matter outside China, where service availability, regulation and app ecosystems differ significantly.
System integration gives HONOR an important advantage
An OS-level agent does not have to interact with every function purely by looking at pixels.
HONOR can expose structured MagicOS tools directly to Qwen Intelligence, providing more reliable access to system features than an ordinary third-party application could receive.
That should reduce some interface-control failures for first-party functions.
Third-party apps remain harder because interfaces change and not every service offers agent-friendly APIs or actions.
Magic9 is still much more than an AI demonstration
HONOR is simultaneously preparing a substantial hardware and camera update for the Magic line.
Official Magic9 material confirms an ARRI collaboration around color science and cinema-oriented imaging modes.
MagicOS 11 also brings new video tools and professional capture features.
Qwen Intelligence should therefore be treated as one major layer of the upcoming flagship rather than the entire Magic9 product story.
The important shift is from answering to acting
The central difference is not that Qwen can produce better conversational responses.
Traditional assistants wait for a relatively precise command. Agent systems attempt to accept a broader objective, work out the path and perform the steps themselves.
That can remove a huge amount of repetitive interaction when the system understands correctly.
It also increases the cost of every misunderstanding.
September 28 should be judged by safeguards, not just demos
A phone completing a hundred-step demonstration is impressive. It does not answer the most important everyday questions.
Will Magic9 require confirmation before payment? Can users restrict specific applications? Is cloud processing clearly disclosed? Can a long task be cancelled and rolled back safely?
Those details are not as exciting on a launch stage as a headline accuracy percentage.
They will determine whether Qwen Intelligence feels like something people can genuinely delegate work to, or an impressive agent they still need to supervise continuously.