We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

thinkingmachines/

Inkling

$1.00

in

$4.05

out

$0.17

cached

/ 1M tokens

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs.

Deploy Private Endpoint
Public
fp8
131,072
JSON
Function
Multimodal
ProjectLicense
thinkingmachines/Inkling cover image
thinkingmachines/Inkling cover image
Inkling

Ask me anything

0.00s

You need to log in to use this model

Log In

Settings

Model Information

Inkling

1. General Information

Inkling is a general-purpose multimodal model that accepts text, image and audio inputs and generates text outputs. It is intended for use in English and other languages, and across multiple coding languages. The model is designed to be used by developers building AI-powered applications, including agentic and tool-use systems, coding assistants, chatbots, and retrieval-augmented generation systems, and is suitable for general-purpose conversational use, instruction-following, and other natural language and multimodal tasks. It is released with open weights to support research, fine-tuning and integration into third-party products by downstream developers.

Languages: English, with general multilingual capabilities across other languages.

2. Model Properties

Model type

Multimodal autoregressive transformer

Architecture type

A 66-layer decoder-only transformer with a sparse Mixture-of-Experts (MoE) feed-forward backbone: each token is routed to 6 of 256 experts, plus 2 shared experts active on every token. Attention is a hybrid of local and global layers. The model is natively multimodal — images and video are encoded via a hierarchical patch encoder, and audio via discrete token encoding — with all modalities projected into a shared hidden space and processed jointly by the decoder.

Parameters

975B total, 41B active

Numerics support

BF16 and NVFP4

Input modalities

Inkling accepts text input in UTF-8 encoding, image input in any pixel-based format (with each dimension ideally between 40px and 4096px for optimal performance), and audio input in WAV format sampled at 16kHz (ideally under 20 minutes in length for optimal performance).

Output modalities

Inkling generates output as UTF-8 encoded text.

3. Evaluations

Inkling results are reported at effort=0.99. Comparison scores are generated Jul 14, 2026. Nemotron 3 Ultra, Kimi K2.5, Kimi K2.6, GLM 5.2, and DeepSeek V4 Pro are open weights models; Gemini 3.1 Pro, Claude Fable 5, and GPT 5.6 Sol are closed weights models.

InklingNemotron 3 UltraKimi K2.5Kimi K2.6GLM 5.2DeepSeek V4 ProGemini 3.1 Pro (high)Claude Fable 5 (max)GPT 5.6 Sol (xhigh)
Reasoning
HLE (text only)29.7%26.6%29.4%35.9%40.1%35.9%44.7%53.3%47.2%
HLE (with tools)46.0%37.4%50.2%54.0%54.7%48.2%51.4%64.5%55.0%
AIME 202697.1%94.2%95.8%96.4%99.2%96.7%98.3%99.9%
GPQA Diamond87.2%86.7%87.9%91.1%89.5%88.8%94.1%92.6%94.1%
Agentic (coding)
SWEBench Verified77.6%70.7%76.8%80.2%80.6%80.6%95.0%
SWEBench Pro (Public)54.3%46.4%50.7%58.6%62.1%55.4%54.2%80.0%64.6%
Terminal Bench 2.1 (Best Harness)63.856.451.371.382.76473.884.689.5
GDPVal-AA v212331164100911901514130796217601748
Agentic (general)
MCP Atlas74.1%44.7%64.0%68.1%77.8%73.2%78.2%83.3%81.8%
Tau 3 Banking23.7%13.8%13.2%20.6%26.8%25.8%16.5%26.8%33.0%
Factuality
BrowseComp (w/ Ctx)77.1%74.9%83.2%83.4%85.9%88.0%89.4%
SimpleQA Verified43.9%32.4%36.9%38.7%38.1%57.0%77.3%68.3%71.6%
AA Omniscience1.0%-1.0%-8.0%6.0%4.0%-10.0%33.0%40.0%22.0%
Chat
IFBench79.8%81.4%70.2%76.0%73.3%76.5%77.1%63.5%72.7%
Global-MMLU-Lite88.7%85.6%84.0%88.4%89.2%89.3%92.7%93.3%91.8%
Vision
MMMU Pro (Standard 10)73.5%75.0%79.0%82.0%84.2%83.0%
Charxiv RQ78.1%77.5%80.4%80.2%86.5%84.7%
Charxiv RQ (with python)82.0%78.7%86.7%89.9%89.4%87.8%
Audio
Audio MC56.6%66.8%
MMAU77.2%82.5%
VoiceBench91.4%94.3%
Safety
FORTRESS (Adversarial)78.0%77.6%54.1%65.6%71.3%36.0%65.2%96.0%82.4%
FORTRESS (Benign)95.9%90.5%98.3%97.2%90.0%98.5%98.0%55.1%98.1%
StrongREJECT98.6%98.7%99.5%99.8%98.5%98.6%98.0%98.7%98.5%