Models

Every model, one app.

Starnexus gives you direct access to the leading models from Xiaomi, Moonshot, Qwen, Zhipu AI, and Google. Pick the right one for each task.

Xiaomi

Efficient open models focused on reasoning and agents, with a small footprint, fast speed, and strong coding.

MiMo UltraSpeed

Plus

The quicker MiMo, for when speed matters more than depth.

Efficient 15B-active Mixture-of-Experts built for high-speed reasoning and agentic workflows. Strong on math and software engineering benchmarks while staying cheap and fast.

MiMo V2.5

Free1M context

Omni-modal agent with 1M context at half the cost.

Xiaomi's native multimodal model. It understands text, images, audio, and video, and delivers Pro-level agentic performance at roughly half the inference cost, with a 1M-token context for long documents and multi-step tasks.

MiMo V2.5 Pro

Plus

Xiaomi's flagship. The cheapest way to run a capable model.

Xiaomi's most capable model, built for complex software engineering and tasks that take humans days. It has a 1M-token context. Xiaomi reports it sustaining 1,000+ tool calls in a single run, and leading ClawEval, GDPVal, and SWE-bench Pro.

Moonshot AI

Specialises in long-context reasoning, agentic workflows, and deep tool-use.

K2.7 Code HighSpeed

Plus

The coding model, served faster.

K2.7 Code on a faster serving tier, for rapid iteration in long coding sessions. Included on the Plus plan.

Kimi K3

Pro

Moonshot's flagship. The most careful model we carry.

Moonshot's most capable model, built for long-horizon coding, agentic workflows with coordinated sub-agents, and deep tool-use over very long contexts. Included on the Pro plan.

Qwen

Qwen builds one of the world's most-used model families, with strong multilingual, coding, and agentic performance.

Qwen3.7 Plus

Free

The balanced middle of the Qwen line.

A cost-effective step up with stronger reasoning and longer context, for research and heavier daily work. Included on the Free plan.

Qwen3.8 Max

Pro1M context

The most capable Qwen model.

A sparse mixture-of-experts model that reads images as well as text, with a 1M-token context for long documents and multi-step work. Alibaba puts it at 2.4 trillion parameters. Included on the Pro plan.

Zhipu AI

Builds the GLM family of general-purpose models, with a dedicated vision line.

GLM 5.3

Pro

Zhipu's flagship. Fast for its size.

Reasons at the level of the largest models while answering far quicker than most of them, at a fraction of the cost. A strong default for hard questions, long documents, and code.

GLM 5V Turbo

Free

Reads images, screens, and documents.

A vision model built to understand images, screenshots, and scanned documents alongside text. Included on the Free plan.

Google

The Gemini family, engineered for multimodal reasoning at scale across text, images, video, audio, and code.

Gemini 3.7 Flash

Plus

A fast, cheap default.

Google's fastest model, tuned for instant answers rather than deep thought. Handles everyday tasks and complex agentic workflows with multimodal input. A strong default for most chats.

Gemini 3.7 Flash Thinking

Pro

Google's deepest reasoning, built for hard problems.

Extended thinking mode over the same multimodal core as the rest of the Gemini family: text, images, video, audio, and code, with a 1M-token context. Optimized for software engineering, agentic workflows, and multi-step reasoning.