NEWS
Arm Bets Compute Subsystems Will Reach Physical AI
Arm is selling compute subsystems as its next royalty unit, from CSS for Mobile 2 to an 80-company physical AI program.
Arm Holdings on September 8 unveiled CSS for Mobile 2, a full compute subsystem aimed at on-phone agents and neural graphics. The drop came at Arm Everywhere China in Shanghai, and Arm shares were at $257.39 in premarket trade, up 2.10% from the September 4 close of $252.09.
The same keynote also pushed Neoverse CSS N4 for cloud chips and Arm Total Design for Physical AI. The through-line is the wager: sell a finished subsystem, collect a fatter royalty, and let phones, servers, and robots share one software path.
A Three-Platform Pitch in Shanghai
Chris Bergey, Arm’s executive vice president of Edge AI, introduced a new AI-native compute platform that packages the C2 CPU cluster, the Mali G2-Ultra NX GPU, system IP, physical layout, and software. Drew Henry, who runs Physical AI, used the same day to expand Total Design into robots and vehicles. Cloud AI leadership, in parallel, put Neoverse CSS N4 in front of customers who already build custom Arm servers.
Arm is not launching a phone chip of its own here. It is selling a recipe. Partners still add their own neural engines, modems, and cameras. The bet is that more of them will take the recipe whole, because stitching CPU, GPU, interconnect, and software eats months they no longer have.
THREE BETS IN ONE KEYNOTE
| Platform | Where it sits | What Arm is selling |
|---|---|---|
| CSS for Mobile 2 | Phones, tablets, other edge devices | C2 CPU cluster, Mali G2-Ultra NX, SI L2 interconnect, software |
| Neoverse CSS N4 | Cloud, networking, DPUs | Up to 128 cores per die, LPDDR6, PCIe Gen 7, pre-validated system IP |
| Total Design for Physical AI | Robots, vehicles, industrial machines | Partner program plus a Robotics Capability Framework |
That split matches how Arm already makes money. Cores still license. Subsystems pull more of the die under Arm’s roof, which is the point of CSS in phones and in the data center.
The Phone Stack Now Runs Agents and Frames
Bergey’s note treats the phone less as an app launcher and more as a device that has to hold context, call tools, and keep doing work inside a thermal budget. The C2 cluster pairs C2-Ultra, Arm’s fastest mobile CPU, with efficiency-focused C2-Pro cores and two SME2 matrix units. That doubles SME2 versus the prior CSS setup, so small language models can stay on the CPU instead of hopping to a separate engine for every token.
Stefan Rosinger, head of product for Edge AI, put a number on the whole loop. Across speech, memory retrieval, reasoning, an app launch, and a browse, C2-Ultra with two SME2 units finishes the job 24% faster than the previous generation. Isolated model scores look bigger. The 24% figure is the one that includes the handoffs, which is where agent jobs actually stall.
C2 CLUSTER AND MALI G2-ULTRA NX
- AI models: C2-Ultra reaches up to 1.7x the prior C1-Ultra on the latest models, with up to 70% speedup on small language models from the extra SME2 block.
- Everyday CPU: Up to 15% higher single-thread speed, 15% faster web browsing, 12% faster app launch, and 12% higher cluster multi-thread performance.
- Power: Up to 38% less power at the same performance versus C1-Ultra.
- Neural graphics: Mali G2-Ultra NX puts neural accelerators inside the shader cores and claims up to 4x performance per watt for neural graphics, plus 14% higher speed on existing game content and a 70% cut in ray-tracing work.
The GPU is the other half of the pitch. Neural Super Sampling, frame-rate upscaling, and denoising run where the pixels already live, rather than shuttling frames to an NPU and back. Arm showed that path in Neural Dawn with Sumo Digital, and it named Tencent Games’ MagicDawn, Unity China’s Tuanjie Engine, NetEase’s Where Winds Meet, Tencent’s Arena Breakout Infinite, and Infold Games’ Infinity Nikki as early software hooks.
CSS for Mobile 2 does not include an NPU. If a partner wants one, it attaches its own. Arm is selling the idea that the CPU keeps state and the GPU keeps frames, with the SI L2 interconnect policing latency, bandwidth, and quality of service while those engines fight over the same memory.
Introducing Arm CSS for Mobile 2. 📱
Built for the system-level demands of agentic AI and cinematic mobile graphics, our new AI-native compute platform brings together world-class AI CPU and GPU capabilities to change what the next generation of mobile devices can do.… pic.twitter.com/jtQmF3OBpj
— Arm (@Arm) September 8, 2026
Software is the lock that makes a subsystem sticky. Sharbani Roy, vice president of AI and developer platforms, launched Arm AI Portal against a base of 22 million developers, with pre-tuned Alibaba Qwen, Google Gemma, and Ultralytics YOLO models on runtimes that include ExecuTorch on-device runtime, LiteRT, and ONNX-RT. On a vivo X300, Arm said Qwen3-TTS ran over 4x faster with SME2, and YOLO26n gained over 40% in a single-thread FP16 path.
Xiaomi Put the New GPU in a Foldable First
The GPU is not waiting on the full CSS. Spec sheets for Xiaomi’s 18 Fold, which launched in China on September 7, list an Arm G2-Ultra NX MC16 GPU on Xiaomi’s in-house Xring O3 chip. August teardown notes still placed last-generation C1 CPU cores on that die. Xiaomi also added its own NPU.
That mix is allowed. Rosinger’s brief says partners can take pieces on their own or combine them with custom and third-party IP. Xiaomi did both, a day before the Shanghai keynote. The new graphics block is in a shipping foldable. The new C2 cluster is not.
SME2 is already in the field in another form. Bergey said it ships in leading Android and iOS handsets, with Alipay, Google AI Edge Gallery, OPPO, and vivo in the software mix. The C2 story is a wider cluster and a second SME2 unit, not the first time matrix math has sat on an Arm CPU.
Together with Arm, we are bringing Arm Neural Technology to the vivo’s latest flagship smartphones, built on Arm’s latest compute platform, to unlock new possibilities for mobile gaming. By combining Arm’s latest compute and graphics technologies with vivo’s device-level optimizations, we aim to accelerate neural graphics adoption and deliver more immersive gaming experiences on mobile.
vivo, statement on Arm’s CSS for Mobile 2 brief
Generated frames will be the sore point in games, the same way they are on PCs. Super sampling can raise a native image. Frame generation can double a counter while adding delay. Arm’s own demos lean on reconstruction first. The 4x neural-graphics efficiency number is a lab ceiling, not a promise that every title becomes four times as fast.
128 Cores Per Die for Agentic Cloud Chips
Neoverse CSS N4 is the cloud twin of the same idea. Arm described a semi-custom platform with eight to 128 Neoverse N4 cores on one die, clocks up to 3.8 GHz, up to 256 MB of L3 cache, DDR5 or LPDDR6, and PCIe Gen 7, with room to scale across chiplets and sockets. Versus Neoverse CSS N3, Arm claims 2x socket performance, 1.25x performance per watt, and 1.75x memory bandwidth. The reference build Arm put forward uses the TSMC N3P process.
WHAT N4 CHANGES IN THE RACK
- Density: Up to 128 cores per die, double the 64-core cap on Neoverse CSS N2.
- Memory and I/O: LPDDR6 and PCIe Gen 7, aimed at agent traffic that moves data as much as it multiplies matrices.
- Role: Custom CPUs, DPUs, networking, and storage chips, the jobs where Arm already sits beside accelerators rather than replacing them.
Rene Haas, Arm’s chief executive, said on the fiscal 2026 earnings circuit that data-center royalties had more than doubled year on year and that he expected another doubling. Arm also told investors its CPU share among top hyperscalers is about 50%, and that more than 1.25 billion Neoverse cores have shipped into data centers. CSS is how those cores get sold as a block instead of a catalog of parts.
Agentic jobs are the excuse for a fatter CPU. A model that calls tools, hits a database, and coordinates other agents spends a lot of time off the GPU. Arm wants that control plane on Neoverse, then on the same architecture at the edge, then on the robot.
Unitree and NXP Sign Onto a Robot Ladder
Henry’s physical-AI note is the longest-dated piece of the bet. Arm estimates those industries hold a $200 billion annual compute opportunity in the 2030s. It is an Arm figure, and it sits well past the current phone cycle. The near-term move is organisational: copy the Total Design program that already exists for cloud chips, and point it at machines that sense, plan, and move.
Arm said the new group already includes more than 80 companies. The named list spans AWS, ECARX, Hugging Face, Liquid AI, NXP, PlusAI, PSYONIC, QNX, Qwen, Siemens, and Unitree Robotics. The first joint project is a Robotics Capability Framework, which Arm’s chief architect Richard Grisenthwaite compared, in a manifesto Arm published with the news, to SAE levels for driving automation.
WHAT THE FRAMEWORK IS FOR
- Shared language: Levels from reactive systems up through context-aware, cognitive, and self-improving machines.
- System needs: Latency, where compute sits, memory and power limits, determinism, and safety, tied to real tasks rather than a single benchmark.
- Who is shaping it: Anaxi Labs, ANYbotics, FMC³ Robotics, Fourier, GALBOT, Gravis Robotics, Lenovo, McKinsey, and Robotec.ai are on the working group Arm named.
The auto version of this play is already in motion. Arm, AWS, Google, HERE, RemotiveLabs, and Siemens built a digital-cockpit reference on Arm Zena CSS so software can be written before the silicon exists. Investor slides have put Zena-style physical-AI royalties on a 2028 start, with Armv9 and CSS still a thin slice of that mix. The robot ladder is how Arm tries to make a fragmented hardware market look like the mobile market it already owns.
What CSS Royalties Pay That Cores Do Not
The cash argument is simpler than the product names. A core license is a slice of a chip. A compute subsystem is CPU, interconnect, memory system, and often GPU, pre-integrated and pre-validated. Arm’s own filings for the quarter ended March 31, 2026 showed how that mix is already moving.
Q4 FYE26 SNAPSHOT
- Total revenue: $1,490 million, up 20% year on year, a record quarter.
- Licensing: $819 million, up 29%.
- Royalties: $671 million, up 11%, with growth Arm listed across smartphones, edge AI, physical AI, and cloud AI.
- Data center: Royalty revenue more than doubled year on year.
By early 2026 Arm had 21 CSS licenses across 12 companies, with five customers shipping chips and the top four Android phone brands already on CSS. The following quarter it signed two more next-generation CSS licenses, one for phones and one for data-center networking chips. Chief financial officer Jason Child said CSS had gone from just under 10% of royalties to well into double digits, and that it could reach upwards of 50% over the next couple of years.
That mix shift is the wager hiding under the Shanghai slides. Phone units do not have to boom if the royalty per chip climbs. Cloud cores do not have to beat x86 on every socket if each Arm socket has more cores at a higher rate. Robots do not have to hit Henry’s 2030s number this year if the software and the CSS habit are in place when the units finally move.
The Next Flagships Will Show If CSS Sticks
CSS for Mobile 2 still has to land in a volume SoC that uses the C2 cluster, not only the new GPU. Xiaomi’s 18 Fold shows a partner can ship G2-Ultra NX with last-gen CPU cores and a house NPU. MediaTek, Qualcomm, Samsung, and the other Android houses can do the same. Arm says that flexibility is a feature. It is also the leak in the royalty story, because a cherry-picked GPU does not pay like a full subsystem.
Neural graphics will be judged in games people already play, on a watt budget a foldable can hold. Agent features will be judged on whether a phone can run a multi-step job without dropping the thread or the battery. Cloud N4 will be judged on designs that take a year or more to tape out. Physical AI will be judged on whether Unitree, NXP, Siemens, and the rest actually share a capability ladder, or whether the 80-firm list stays a keynote slide.
Arm put the phone, the rack, and the robot on one stage in Shanghai because it wants them on one architecture and one contract shape. The GPU is already in a foldable. The fatter royalty arrives only if the rest of the subsystem follows it into silicon.
Disclaimer: This article is news reporting and analysis of Arm’s September 8 product announcements and related financial figures. It is for information only and is not investment advice, a recommendation to buy or sell Arm Holdings ADRs, or a forecast of future royalties, shipments, or partner adoption. Readers who are considering a position in semiconductor or robotics-linked stocks should consult a licensed financial adviser who can review their own holdings, time horizon, and risk limits. Revenue, royalty, stock-price, and partner figures reflect the company filings, newsroom notes, and market prints cited here and can change with later earnings, design wins, or trading sessions.
-
BUSINESS4 days agoOstrich Farms Rebuild Around a Tick-Bite Meat Allergy
-
BUSINESS3 days agoThe Yen Rally Was Funded by a Record Reserve Sale
-
BUSINESS1 week agoOpenAI Turns ChatGPT Into a $1 Billion Ad Auction
-
NEWS2 weeks agoThe Newest Chip Is in India’s Cheapest Mac Mini
-
NEWS3 days agoCoremail Pitches AI-Native Email Security at LEAP 2026
-
NEWS2 days agoGottheimer’s AI LABS Act Enters an Already Crowded Field
