We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

thinkingmachines/

Inkling-Small

$0.58

in

$1.44

out

$0.116

cached

/ 1M tokens

Inkling-Small is a Mixture-of-Experts transformer with 276B total parameters, 12B active, trained on NVIDIA GB300 NVL72 systems. Like Inkling, it features native reasoning over audio and images, variable thinking effort

Deploy Private Endpoint
Public
fp8
524,288
JSON
Function
Multimodal
ProjectPaperLicense
thinkingmachines/Inkling-Small cover image
thinkingmachines/Inkling-Small cover image
Inkling-Small

Ask me anything

0.00s

You need to log in to use this model

Log In

Settings

Model Information

license: apache-2.0 license_link: https://www.apache.org/licenses/LICENSE-2.0 pipeline_tag: image-text-to-text tags:

  • conversational
  • image-text-to-text
  • audio-text-to-text
  • moe library_name: transformers

Inkling

1. General Information

Inkling-Small is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

Languages: English, with general multilingual capabilities across other languages.

2. Model Properties

Model type

Multimodal autoregressive transformer

Architecture type

A 42-layer decoder-only transformer with a sparse Mixture-of-Experts (MoE) feed-forward backbone: each token is routed to 6 of 256 experts, plus 2 shared experts active on every token. Attention is a hybrid of local and global layers. The model is natively multimodal — images are encoded via a hierarchical patch encoder, and audio via discrete token encoding — with all modalities projected into a shared hidden space and processed jointly by the decoder.

Parameters

276B total, 12B active

Numerics support

BF16 and NVFP4

Input modalities

Inkling-Small accepts text, image, and audio inputs:

  • Text: UTF-8 encoded text
  • Image: Any pixel-based image input. For optimal performance, each image dimension should be between 40px to 4096px.
  • Audio: WAV format, sampled at 16kHz. For optimal performance, audio length should ideally be under 2 mins.

Output modalities

Inkling-Small generates output as UTF-8 encoded text.

3. Evaluations

Open weights Closed weights
Inkling-Small Qwen3.5 397B-A17B MiMo V2.5 Minimax M2.7 DeepSeek V4 Flash Nemotron 3 Ultra Inkling Claude 4.5 Haiku Gemini 3.5 Flash-Lite GPT 5.6 Luna
Model Info
AA Index (v4.1) 40.0%34.0%37.0%38.0%40.0%38.0%41.0%30.0%36.0%49.0%
Params (B) (activated / total) 12 / 27617 / 39715 / 31010 / 23013 / 28455 / 55041 / 975
Agentic (coding)
SWEBench Verified 80.2%76.4%71.0%79.9%79.0%70.7%77.6%73.3%75.0%93.0%
SWEBench Pro (public) 55.9%50.9%56.1%56.2%52.6%46.4%54.3%39.5%54.2%62.7%
Terminal Bench 2.1 (best harness) 64.7%51.3%63.7%55.4%61.8%56.4%63.8%44.2%54.0%82.5%
SciCode 48.7%42.0%43.1%47.0%44.9%39.9%46.1%43.3%40.9%50.0%
Agentic (general)
GDPval-AA v2 12699621145115911891164123891111391530
MCP Atlas (public / all) 79.6/79.2%74.2%/–49.4%/–69.0%/–47.4/44.7%78.8/76.0%41.2/40.2%79.8/76.8%77.0/75.0%
Tau 3 Banking 15.5%13.4%6.6%8.9%22.9%13.8%23.7%9.1%16.5%24.3%
BrowseComp (with context management) 77.4%78.6%76.3%73.2%63.0%77.1%84.0%
Toolathlon Verified 54.4%40.7%49.1%47.5%50.9%34.3%45.5%26.9%57.1%67.9%
AA-Briefcase 917833870839612
Reasoning (general)
GPQA Diamond 89.5%89.3%84.9%87.4%89.4%86.7%87.2%67.2%83.8%89.5%
HLE (text only) 31.6%27.3%25.2%28.1%32.1%26.6%29.7%9.7%17.5%35.6%
HLE (with tools) 47.8%48.3%40.0%40.3%45.1%37.4%46.0%17.8%42.5%48.9%
AIME 2026 95.5%93.3%93.6%87.7%95.8%94.2%97.1%85.1%82.2%97.6%
HMMT Feb 2026 90.2%87.9%82.6%71.2%93.9%78.8%86.3%66.7%63.6%98.5%
CritPt 8.3%1.7%3.7%0.6%7.1%3.1%5.4%0.0%0.0%20.6%
Reasoning (abstract)
ARC-AGI-1 84.0%79.5%47.7%87.7%
ARC-AGI-2 40.1%36.5%4.0%47.6%
Factuality
SimpleQA Verified 20.6%26.0%16.1%13.5%34.1%32.4%43.9%5.9%44.1%41.7%
AA Omniscience (index) -9.0-29.8-9.30.7-22.9-1.02.1-4.26.9-11.6
Chat
IFBench 82.2%78.8%67.1%75.7%79.2%81.4%79.8%54.3%78.6%67.3%
Global-MMLU-Lite 86.7%90.0%83.5%83.9%88.4%85.6%88.7%83.4%89.4%88.7%
Safety
StrongREJECT 98.4%99.4%99.3%99.4%97.4%98.7%98.6%98.6%97.6%98.7%
FORTRESS (adversarial) 71.6%77.3%64.8%86.3%32.0%77.6%78.0%91.3%70.7%83.8%
FORTRESS (benign) 96.9%95.4%94.6%90.1%99.2%90.6%95.9%94.1%95.5%97.8%
Vision
MMMU Pro (Standard 10) 74.0%77.3%75.4%73.5%58.6%79.0%78.6%
Charxiv RQ (original / with python) 77.4/81.3%80.8%/–81.0%/–78.1/82.0%57.4%/–70.0%/–81.4%/–
Audio
Audio MC 54.9%30.4%56.6%33.6%
MMAU 77.0%73.6%77.2%75.2%
VoiceBench 90.1%86.4%91.4%85.9%
  • Inkling-Small against open-and closed-weights models across the full eval suite. Activated and total parameters are given for scale; a dash means the score was not available at the time of writing.
  • SWEBench Verified: Inkling and Inkling-Small’s numbers are reported using a bash-only harness. We use self-reported numbers for external models.
  • Terminal Bench 2.1: Inkling and Inkling-Small’s numbers are reported using an internal coding harness. A small number of solutions were found to be contaminated from web search and were assigned a score of 0. We use self-reported numbers for external models where available. Otherwise, we report performance using our internal harness.
  • Audio MC: Other models were evaluated internally since they are not on the official leaderboard.
  • VoiceBench: VoiceBench uses rule-based, hard-coded string matching for grading, making the evaluation sensitive to output-formatting differences. We therefore added a system message instructing models to follow the expected answer format.
  • HLE with tools: We benchmarked Minimax M2.7, Claude 4.5 Haiku, Gemini 3.5 Flash-Lite, and GPT 5.6 Luna using our internal harness.