Model catalog

Models ready for agentic tools.

Use these model IDs with any OpenAI-compatible SDK or coding agent pointed at the Aegis proxy. The catalog below lists the frontier models Aegis exposes, including context limits, capabilities, and token pricing.

Models

12

Billing

USD usage

API shape

OpenAI compatible

Available Models

Model IDs are the exact strings to place in the `model` field of chat completion requests.

Reference energy uses 0.35 J per input token

Qwen

Qwen3.5 397B

Reasoning

Qwen3.5 397B-A17B — quant-agnostic canonical name for the Qwen3.5 397B family.

Context

262.13K

Input

$0.6900/M

Energy

~1.75 Wh

ReasoningToolsJSON

Output

$4.14/M

Cached input

$0.1725/M

Images

None

Model ID

qwen3.5-397b
Reference request cost$0.022

Based on a 18K-token reference input before Aegis optimization.

Qwen

Qwen3.5 397B Fast

Fast

Qwen3.5 397B with thinking mode disabled for faster, lower-latency responses. Ideal for straightforward tasks where extended reasoning is unnecessary.

Context

262.13K

Input

$0.6900/M

Energy

~1.17 Wh

ToolsJSON

Output

$4.14/M

Cached input

$0.1725/M

Images

None

Model ID

qwen3.5-397b-fast
Reference request cost$0.015

Based on a 12K-token reference input before Aegis optimization.

MoonshotAI

Kimi K2.6

Reasoning

Kimi K2.6 — quant-agnostic canonical name for the Kimi K2.6 family.

Context

262.13K

Input

$0.6900/M

Energy

~1.75 Wh

ReasoningVisionToolsJSON

Output

$3.22/M

Cached input

$0.1725/M

Images

20

Model ID

kimi-k2.6
Reference request cost$0.022

Based on a 18K-token reference input before Aegis optimization.

MoonshotAI

Kimi K2.6 Fast

Fast

Moonshot Kimi K2.6 with thinking mode disabled for instant, lower-latency responses.

Context

262.13K

Input

$0.6900/M

Energy

~1.17 Wh

VisionToolsJSON

Output

$3.22/M

Cached input

$0.1725/M

Images

20

Model ID

kimi-k2.6-fast
Reference request cost$0.015

Based on a 12K-token reference input before Aegis optimization.

Qwen

Qwen3.6 35B

Reasoning

Qwen3.6 35B-A3B — quant-agnostic canonical name for the Qwen3.6 35B family.

Context

131.06K

Input

$0.2900/M

Energy

~1.56 Wh

ReasoningVisionToolsJSON

Output

$1.15/M

Cached input

$0.0725/M

Images

4

Model ID

qwen3.6-35b
Reference request cost$0.019

Based on a 16K-token reference input before Aegis optimization.

Qwen

Qwen3.6 35B Fast

Fast

Qwen3.6 35B with thinking mode disabled for instant responses. A lightweight, fast model for rapid iteration.

Context

131.06K

Input

$0.2900/M

Energy

~0.97 Wh

VisionToolsJSON

Output

$1.15/M

Cached input

$0.0725/M

Images

4

Model ID

qwen3.6-35b-fast
Reference request cost$0.012

Based on a 10K-token reference input before Aegis optimization.

MoonshotAI

Kimi K2.7 Code

Reasoning

Kimi K2.7 Code — quant-agnostic canonical name for the Kimi K2.7 Code family.

Context

262.13K

Input

$0.9500/M

Energy

~1.75 Wh

ReasoningVisionToolsJSON

Output

$4.00/M

Cached input

$0.2375/M

Images

20

Model ID

kimi-k2.7-code
Reference request cost$0.022

Based on a 18K-token reference input before Aegis optimization.

MoonshotAI

Kimi K3

Reasoning

Moonshot AI Kimi K3 — multimodal 2.8T MoE with a 1M-token context window.

Context

1.05M

Input

$3.00/M

Energy

~1.75 Wh

ReasoningVisionToolsJSON

Output

$15.00/M

Cached input

$0.3000/M

Images

20

Model ID

kimi-k3
Reference request cost$0.022

Based on a 18K-token reference input before Aegis optimization.

ZhipuAI

GLM-5.2

Reasoning

Private GLM-5.2 test canary

Context

1.05M

Input

$1.45/M

Energy

~1.75 Wh

ReasoningTools

Output

$4.50/M

Cached input

$0.3625/M

Images

None

Model ID

glm-5.2
Reference request cost$0.022

Based on a 18K-token reference input before Aegis optimization.

ZhipuAI

GLM-5.2 (fast)

Fast

GLM-5.2 with thinking skipped (reasoning_effort=none). Non-thinking tier, mirrors glm-5.1-fast.

Context

1.05M

Input

$1.45/M

Energy

~1.17 Wh

Tools

Output

$4.50/M

Cached input

$0.3625/M

Images

None

Model ID

glm-5.2-fast
Reference request cost$0.015

Based on a 12K-token reference input before Aegis optimization.

ZhipuAI

GLM-5.2 (short)

Reasoning

GLM-5.2 with a 200K context window and a bounded reasoning budget — faster, lower-energy serving for everyday coding, agentic, and chat workloads under 200K tokens. Private preview (grant-gated).

Context

199.98K

Input

$1.45/M

Energy

~1.36 Wh

ReasoningTools

Output

$4.50/M

Cached input

$0.3625/M

Images

None

Model ID

glm-5.2-short
Reference request cost$0.017

Based on a 14K-token reference input before Aegis optimization.

ZhipuAI

GLM-5.2 (short, fast)

Fast

GLM-5.2 short with reasoning OFF — the fastest, lowest-energy variant of the 200K pool, for latency-sensitive coding, chat, and tool-use that does not need chain-of-thought. Private preview (grant-gated).

Context

199.98K

Input

$1.45/M

Energy

~0.97 Wh

Tools

Output

$4.50/M

Cached input

$0.3625/M

Images

None

Model ID

glm-5.2-short-fast
Reference request cost$0.012

Based on a 10K-token reference input before Aegis optimization.

Integration

Paste the ID into your agent config.

Aegis keeps the endpoint OpenAI-compatible. Keep your Aegis key in the API key field, set the base URL once, then choose a model ID from the table for each agent profile.

Model IDProviderContextInput priceReference energy
qwen3.5-397bQwen262.13K$0.6900/M~1.75 Wh
qwen3.5-397b-fastQwen262.13K$0.6900/M~1.17 Wh
kimi-k2.6MoonshotAI262.13K$0.6900/M~1.75 Wh
kimi-k2.6-fastMoonshotAI262.13K$0.6900/M~1.17 Wh
qwen3.6-35bQwen131.06K$0.2900/M~1.56 Wh
qwen3.6-35b-fastQwen131.06K$0.2900/M~0.97 Wh
kimi-k2.7-codeMoonshotAI262.13K$0.9500/M~1.75 Wh
kimi-k3MoonshotAI1.05M$3.00/M~1.75 Wh
glm-5.2ZhipuAI1.05M$1.45/M~1.75 Wh
glm-5.2-fastZhipuAI1.05M$1.45/M~1.17 Wh
glm-5.2-shortZhipuAI199.98K$1.45/M~1.36 Wh
glm-5.2-short-fastZhipuAI199.98K$1.45/M~0.97 Wh