Project Zenith brings local 30B AI to Windows, but dense models remain a bottleneck
MoE models reach up to 100 tokens per second on Zenith hardware, while dense 70B models manage about five

Microsoft officially named its Windows developer configuration Project Zenith on September 4, 2026, and set a minimum hardware requirement. The company says the platform can run AI models with more than 30 billion parameters locally, without metered cloud charges, according to Logan Iyer, Microsoft's corporate vice president for Windows Platform and Developer.
But the 30B-plus figure describes memory capacity, not performance. Whether a model runs at a useful speed depends largely on whether it is a dense model or a sparse mixture-of-experts (MoE) model.
The first qualifying device, Lenovo's 1.6-litre ThinkCentre X Ultra compact desktop, will go on sale in November 2026 from US$3,699. The machine is manufactured by a company subject to China's National Intelligence Law, a factor that organisations in Myanmar and elsewhere should consider before deploying it for proprietary code, AI model weights or other sensitive material.
Model architecture matters more than the 30B headline
Project Zenith requires at least 64GB of unified memory and 250GB per second of memory bandwidth. Unified memory allows the CPU and GPU to share one pool of RAM, making it possible to load a 30-billion-parameter model on hardware priced closer to a high-end laptop than an enterprise server.
However, loading a model is not the same as running it quickly. Inference on these systems is primarily limited by memory bandwidth: the machine can generate tokens only as quickly as it can move model weights from memory to its processing units.
The AMD Ryzen AI Max+ 395 platform at the centre of the first Zenith systems offers about 256GB/s of memory bandwidth. Apple's Mac Studio M3 Ultra offers roughly 800GB/s. Independent benchmark data indicates that a dense 70-billion-parameter model running at 4-bit quantisation produces about five tokens per second on the Ryzen AI Max+ 395.
That speed is unlikely to be practical for most coding assistants. MoE models perform considerably better because they store all their parameters in memory but activate only a small group of specialised sub-networks for each token.
A 30B MoE model with about 3 billion active parameters per token can produce roughly 70–100 tokens per second on the same class of hardware. Benchmarks have measured Qwen3-30B-A3B at about 34 tokens per second and GPT-OSS 120B at about 39 tokens per second on comparable systems.
The practical lesson is that developers should check performance for the exact model they plan to use. Dense models, including some Llama, Gemma and Mistral variants, may deliver fewer than 10 tokens per second. Qwen3, Mixtral-class and similar MoE models are more likely to provide a responsive local experience. Microsoft's announcement did not identify either model category or a token-generation speed.
What Project Zenith actually is
Project Zenith is not a new edition of Windows. It is a factory-applied software configuration for hardware meeting Microsoft's minimum specifications.
Devices ship with Visual Studio Code, Windows Terminal, GitHub Copilot, PowerToys, Git, Python 3.14 or later, Node.js 24 or later through NVM, Windows Subsystem for Linux 2 with Ubuntu, .NET 10 and the WinAppCLI toolchain. File Explorer is configured to display file extensions, hidden files and the full path in the title bar. Long-path support is enabled, while Start menu tips, recently used files, synchronisation-provider notifications and account prompts are disabled.
Agent-security features previewed at Build 2026, including operating-system-enforced identity verification and Microsoft Execution Containers for isolating agentic workloads, are enabled from the first boot.
Developers can apply the same configuration to an existing Windows 11 machine using Microsoft's Windows Developer Configuration script on GitHub. The hardware requirement remains: local inference for 30B-class models still requires at least 64GB of unified memory.
Sceptics question the preconfigured experience
Windows analyst Paul Thurrott tested the public configuration and described it as a "curious miscalculation". His argument is that developers already have preferred environments and may spend time undoing or changing Microsoft's settings after buying a machine.
"I had to wipe the PC I tried this on, it was maddening," Thurrott said. Windows Central's Sean Endicott took a more neutral view, noting that tools valued by developers could appear as bloat to general users.
Microsoft says Project Zenith reflects developer feedback and that the configuration will evolve with developers and the wider community.
Why local inference can make economic sense
The case for a US$3,699 workstation is not limited to convenience. Signal65 Research founder Ryan Shrout says autonomous agents run continuously and consume far more tokens than ordinary chat. His analysis estimates usage at roughly four to 15 times that of conversational AI, with further growth expected.
For teams running agentic coding assistants throughout the working day, cloud inference charges can rise quickly under per-token pricing. A local machine could eventually recover its purchase cost by reducing recurring cloud fees, but the calculation depends on the buyer's own usage and billing history.
First hardware: Lenovo ThinkCentre X Ultra
The first Project Zenith systems use AMD's Ryzen AI Halo platform, which combines CPU, GPU and NPU processing in a shared memory pool. AMD unveiled its own Ryzen AI Halo mini-PC at IFA 2026 in Berlin on September 4, but did not disclose pricing.
The first confirmed Project Zenith machine is Lenovo's ThinkCentre X Ultra, announced at IFA on September 3. It uses the AMD Ryzen AI Max+ PRO 495, a 16-core Zen 5 processor with a Radeon 8065S integrated GPU offering about 55 TOPS of neural-processing capability. The system supports up to 128GB of LPDDR5X unified memory and will ship in November 2026 from US$3,699.
Lenovo also says up to four ThinkCentre X Ultra units can be linked, pooling as much as 512GB of memory and about 524 total TOPS of AI compute. The arrangement is intended to support very large models, including Meta's Llama 4 Maverick, without rack-mounted servers. Lenovo has not disclosed the effective interconnect bandwidth, so the performance of distributed inference cannot yet be verified.
Nvidia's DGX Spark offers 273GB/s of bandwidth and is available for US$4,699, up from its October 2025 launch price of US$3,999 after memory-supply constraints led to a price increase in February 2026. Microsoft says more hardware and silicon partners will join the programme, with Nvidia's RTX Spark platform widely expected to be among them.
What independent benchmarks show
Independent testing gives a more detailed picture than vendor announcements. LTT Labs benchmarks reported by Gigazine in July 2026 found that the Mac Studio M3 Ultra generated tokens roughly two to three times faster than Ryzen AI Max+ 395 systems on dense models including Gemma 4. The difference was attributed to Apple's approximately 800GB/s memory bandwidth compared with AMD's 256GB/s.
AMD's own comparisons with Nvidia's DGX Spark show Ryzen AI Halo producing about 7% more tokens per second on GPT-OSS 120B and 12% more on Qwen 3.5 122B, both MoE models. Those figures are vendor claims and require independent verification. Shrout said AMD's MoE results should not be treated as settled until independently validated.
Mac Studio offers greater raw bandwidth and a mature local-inference software stack through MLX and Metal. DGX Spark provides the CUDA ecosystem and stronger prompt-processing performance for prefill-heavy agentic workloads. Project Zenith sits between them: it is more configurable than Apple's platform and less dependent on CUDA than Nvidia's, while offering a factory-prepared developer environment.
Security questions surrounding Lenovo
Lenovo is headquartered in Beijing and its majority shareholder is Legend Holdings, a Chinese entity whose largest shareholder is the Chinese Academy of Sciences, a state institution.
China's National Intelligence Law, enacted in 2017, requires organisations and citizens to support, assist and cooperate with national intelligence efforts in accordance with the law. China's Cybersecurity Law also requires cooperation with security inspections. These are legal conditions that apply to Lenovo as a Chinese organisation, regardless of where a product is sold or of the structure of its overseas subsidiaries.
For developer workstations, relevant areas of concern include:
- Usage telemetry collected by Lenovo Vantage
- AI inference queries captured by telemetry systems
- Network configuration data in enterprise deployments
- Biometric data if fingerprint authentication is used
No independent security audit specifically covering the ThinkCentre X Ultra had been published when the device was announced on September 3, 2026. Buyers should account for that gap before making enterprise procurement decisions.
Lenovo's wider history includes a classified-network ban by intelligence agencies in the United States, the United Kingdom, Australia, Canada and New Zealand in the mid-2000s; a 2015 Superfish adware incident that led to an FTC complaint settled in 2018; and a 2026 class-action lawsuit alleging that Lenovo's advertising technology enabled bulk transfers of sensitive personal identifiers to China under the US Justice Department's Bulk Sensitive Data Transfer Rule. Lenovo has denied the claims and said it takes data security seriously.
Buyers can disable or uninstall Lenovo Vantage, review BIOS telemetry settings and use network segmentation in enterprise deployments. These measures do not remove the underlying legal exposure under Article 7 of China's National Intelligence Law. Organisations handling classified, export-controlled or commercially sensitive material should consult their security teams before buying the system.
Several US states have designated Lenovo a prohibited supplier. A 2023 letter from the US House Select Committee urged the US Navy Exchange to remove Lenovo hardware from military retail outlets.
Why the configuration is open source
Microsoft has published the full Windows Developer Configuration script on GitHub. Developers can use it through winget to apply the same tools and settings to an existing Windows 11 installation. A restart is required when WSL is enabled.
As a result, Project Zenith does not require a new PC in every case. It requires hardware with at least 64GB of unified memory and 250GB/s of bandwidth. Owners of an existing qualifying Ryzen AI Max+ system can use the software configuration now.
The ecosystem is still limited
Project Zenith currently has one confirmed hardware maker, Lenovo, and one confirmed silicon platform, AMD Ryzen AI Halo. Microsoft says more OEM and silicon partners will arrive in the coming months. Nvidia's RTX Spark ecosystem is the most anticipated addition.
The strategic case is straightforward: as AI providers raise per-token prices and coding agents shift from short conversations to continuous workflows, local inference becomes more attractive. Microsoft's strategy is to persuade developers to buy their own inference capacity rather than rent it from the cloud. Whether the strategy succeeds will depend on how quickly AMD, Nvidia and their OEM partners expand the 64GB-plus unified-memory market.
Should developers buy now?
Check performance first: Identify the model family you will use most. MoE systems such as Qwen3, Mixtral variants and GPT-OSS-class models may reach 70–100 tokens per second on qualifying hardware. Dense models such as Llama 3 70B and comparable architectures may produce fewer than 10 tokens per second on 256GB/s hardware.
Check cloud spending next: Review the team's actual cloud-inference costs for agentic coding over the past 30 days. If spending is close to or above US$100 per developer each month, the payback period on a US$3,699 machine becomes measurable. If spending is substantially lower, the purchase may not make economic sense on its own.
Assess security before procurement: Teams handling proprietary code, trade secrets or regulated data should consult their security specialists about Lenovo's ownership and the legal obligations under China's National Intelligence Law. No independent security audit of the ThinkCentre X Ultra is currently available.
Waiting is also an option: If Nvidia's RTX Spark ecosystem joins Project Zenith as expected, hardware choices should broaden. Developers who can wait 30–90 days may have more options.
Frequently Asked Questions
Is Project Zenith a new version of Windows?
No. Project Zenith is a factory-applied software configuration for systems with at least 64GB of unified memory and 250GB/s of memory bandwidth. It installs developer tools including VS Code, GitHub Copilot, Python, Node and WSL, applies developer-oriented Windows settings and enables Microsoft's agentic security features. The same configuration can be installed on an existing Windows 11 machine if the hardware meets the requirements.
Can it run a 30B AI model at useful speeds?
It depends on the architecture. A 30B MoE model that activates only part of its parameters per token can generate roughly 70–100 tokens per second on qualifying hardware. Dense models are considerably slower. A dense 70B model at 4-bit quantisation produces about five tokens per second on the AMD Ryzen AI Max+ 395 hardware used by the first Zenith systems. Developers should check the exact model before treating the platform as a replacement for cloud inference.
What are the privacy and security risks of the Lenovo system?
Lenovo is a Chinese-owned company subject to China's 2017 National Intelligence Law, including Article 7, which requires cooperation with government intelligence efforts. No independent security audit of the ThinkCentre X Ultra was available when it was announced on September 3, 2026. Organisations handling proprietary code, model weights or regulated data should assess the issue with their security teams. Network segmentation and disabling Lenovo telemetry software may reduce practical exposure but do not remove the underlying legal obligation.
How does it compare with the Mac Studio and Nvidia DGX Spark?
The Mac Studio M3 Ultra has about 800GB/s of memory bandwidth, around three times the Ryzen AI Max+ 395's 256GB/s, and independent tests show two to three times faster token generation on dense models. Nvidia's DGX Spark offers the CUDA ecosystem and about 39 tokens per second on GPT-OSS 120B, compared with roughly 34 on AMD in cited benchmarks. Project Zenith offers a more open Windows-and-Linux configuration, with ROCm support, at a starting price of US$3,699 compared with US$4,699 for the DGX Spark as of February 2026. The appropriate choice depends on model architecture, software requirements and security policy rather than price alone.
Originally published on Tech Times
ⓒ {{Year}} TECHTIMES.com All rights reserved. Do not reproduce without permission.





















