Starnexus gives you direct access to the leading models from Xiaomi, Moonshot, Qwen, Zhipu AI, and Google. Pick the right one for each task.
Efficient open models focused on reasoning and agents, with a small footprint, fast speed, and strong coding.
The quicker MiMo, for when speed matters more than depth.
Efficient 15B-active Mixture-of-Experts built for high-speed reasoning and agentic workflows. Strong on math and software engineering benchmarks while staying cheap and fast.
Omni-modal agent with 1M context at half the cost.
Xiaomi's native multimodal model. It understands text, images, audio, and video, and delivers Pro-level agentic performance at roughly half the inference cost, with a 1M-token context for long documents and multi-step tasks.
Xiaomi's flagship. The cheapest way to run a capable model.
Xiaomi's most capable model, built for complex software engineering and tasks that take humans days. It has a 1M-token context. Xiaomi reports it sustaining 1,000+ tool calls in a single run, and leading ClawEval, GDPVal, and SWE-bench Pro.
Specialises in long-context reasoning, agentic workflows, and deep tool-use.
The coding model, served faster.
K2.7 Code on a faster serving tier, for rapid iteration in long coding sessions. Included on the Plus plan.
Moonshot's flagship. The most careful model we carry.
Moonshot's most capable model, built for long-horizon coding, agentic workflows with coordinated sub-agents, and deep tool-use over very long contexts. Included on the Pro plan.
Qwen builds one of the world's most-used model families, with strong multilingual, coding, and agentic performance.
The balanced middle of the Qwen line.
A cost-effective step up with stronger reasoning and longer context, for research and heavier daily work. Included on the Free plan.
The most capable Qwen model.
A sparse mixture-of-experts model that reads images as well as text, with a 1M-token context for long documents and multi-step work. Alibaba puts it at 2.4 trillion parameters. Included on the Pro plan.
Builds the GLM family of general-purpose models, with a dedicated vision line.
Zhipu's flagship. Fast for its size.
Reasons at the level of the largest models while answering far quicker than most of them, at a fraction of the cost. A strong default for hard questions, long documents, and code.
Reads images, screens, and documents.
A vision model built to understand images, screenshots, and scanned documents alongside text. Included on the Free plan.
The Gemini family, engineered for multimodal reasoning at scale across text, images, video, audio, and code.
A fast, cheap default.
Google's fastest model, tuned for instant answers rather than deep thought. Handles everyday tasks and complex agentic workflows with multimodal input. A strong default for most chats.
Google's deepest reasoning, built for hard problems.
Extended thinking mode over the same multimodal core as the rest of the Gemini family: text, images, video, audio, and code, with a 1M-token context. Optimized for software engineering, agentic workflows, and multi-step reasoning.