We use essential cookies to make our site work. With your consent, we may also use non-essential cookies to improve user experience and analyze website traffic…

DeepInfra raises $107M Series B to scale the inference cloud — read the announcement

Audio8/

Audio8-TTS-Preview-0.6b

$5.00

/ 1M characters

A 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning, supporting 11 languages.

Audio8/Audio8-TTS-Preview-0.6b cover image

Input

Input text

Text to convert to speech

You need to log in to use this model

Log In

Settings

ServiceTier

The service tier used for processing the request. 'priority' processes the request with higher priority (premium rate); 'flex' processes it at lower priority for a discount, served only when spare capacity exists and may be retried/timed out under load. Both apply only to models that support the respective tier. For compatibility, 'auto' is treated as 'priority' and 'standard_only' as 'default'.

Fail Fast

If true, the request is rejected immediately with HTTP 429 when the model has no spare capacity, instead of waiting in the queue. Opt-in; the default (false) keeps standard queueing behavior.

TtsResponseFormat

Select the desired format for the speech output. Supported formats include mp3, opus, flac, wav, and pcm.

Max new tokens

Controls the maximum length of the generated audio per chunk (more tokens = longer audio). (Default: 1024, 32 ≤ max_new_tokens ≤ 2048)

Temperature

Sampling temperature; lower is more deterministic. (Default: 0.7, 0 ≤ temperature ≤ 2)

Top P

Nucleus sampling probability mass. (Default: 0.9, 0 ≤ top_p ≤ 1)

Top K

Restrict sampling to the top K tokens. (Default: 50, 0 ≤ top_k ≤ 1000)

Seed

Seed for the random number generator. (Default: empty, 0 ≤ seed ≤ 2147483647)

Output

Waiting for audio data... Submit request to start streaming.

Model Information
Audio8

  Audio8 TTS Preview 0.6B: SOTA-Class TTS at Compact Scale

A 0.6B-parameter multilingual text-to-speech model with zero-shot voice cloning.

GitHub Demo ONNX INT4 License

Audio8 TTS Preview supports multilingual speech generation and zero-shot voice cloning. This repository contains the complete checkpoint, its 44.1 kHz neural audio codec, tokenizer, processor, and Hugging Face remote code.

Preview status: Language coverage is intentionally limited in this release. For the best results, use one of the 11 recommended languages below. Broader multilingual coverage and Chinese dialect support are planned for future releases.

Supported Languages

Cantonese  ·  Chinese  ·  Dutch  ·  English
French  ·  German  ·  Italian  ·  Japanese
Korean  ·  Polish  ·  Spanish

Model Details

Audio8 TTS uses a DualAR architecture inspired by Fish Audio S2 Pro. The slow AR transformer predicts one semantic token for each audio frame. The fast AR transformer predicts the frame's codec codebooks, conditioned on the slow hidden state and preceding codebooks.

ComponentConfiguration
Main model601,159,424 parameters, excluding the codec
Slow AR24 layers, width 896, 14 attention heads, 2 KV heads
Fast AR4 layers, width 896, 14 attention heads, 2 KV heads
Acoustic tokens10 codebooks, 4,096 entries per codebook
Codec44.1 kHz, 2,048 samples per model frame (~21.5 frames/s)
ContextUp to 2,048 packed text/audio positions

The bundled codec handles both reference-audio encoding and waveform decoding, so no additional codec checkpoint is required.

Installation

Python 3.10 or newer and a CUDA-capable GPU are recommended.

pip install "torch>=2.5.0" "torchaudio>=2.5.0" \
  "transformers>=4.57.0,<5" "soundfile>=0.12" "safetensors>=0.4"
copy

Usage

The model uses custom Transformers code. Review the files in this repository, then load it with trust_remote_code=True.

Zero-shot voice cloning

The reference transcript must match the spoken content in the reference audio.

import soundfile as sf
import torch
from transformers import AutoModel, AutoProcessor

model_id = "AutoArk-AI/Audio8-TTS-Preview-0.6b"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32

processor = AutoProcessor.from_pretrained(model_id, trust_remote_code=True)
model = AutoModel.from_pretrained(
    model_id,
    trust_remote_code=True,
    dtype=dtype,
).eval().to(device)

inputs = processor(
    text=["Welcome to Audio8 TTS."],
    reference_audio=["reference.wav"],
    reference_text=["The exact transcript of the reference recording."],
    return_tensors="pt",
)
inputs = {name: value.to(device) for name, value in inputs.items()}

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=1024,
        temperature=0.8,
        top_p=0.95,
        top_k=50,
        do_sample=True,
        return_dict_in_generate=True,
    )
    waveforms, waveform_lengths = model.decode_audio(output.codes)

audio = waveforms[0, : int(waveform_lengths[0])].float().cpu().numpy()
sf.write("output.wav", audio, model.config.codec_sample_rate)
copy

Generation without a reference

Omit reference_audio and reference_text when a cloned voice is not needed:

inputs = processor(
    text=["This utterance does not use a reference voice."],
    return_tensors="pt",
)
copy

For command-line inference, batching, and supervised fine-tuning, see the Audio8 TTS repository.

Deployment Options

CPU deployment: ONNX INT4

Audio8-TTS-Preview-0.6B-ONNX-INT4 packages Audio8 TTS for low-resource CPU inference with ONNX Runtime. Slow and Fast AR weights use weight-only INT4, while activations, KV caches, and the neural audio codec use FP16.

AdvantageDetails
CPU nativeRuns with ONNX Runtime CPUExecutionProvider; no CUDA required
Low memoryAbout 1 GiB after loading in the tested Apple M2 configuration
Small runtimeNo PyTorch or Transformers dependency after model download
Complete workflowCLI, web and HTTP service, streaming PCM, and voice registration

Normal synthesis loads only the Slow AR, Fast AR, and codec decoder sessions. Voice registration releases those sessions before loading the optional codec encoder, keeping peak memory controlled.

Get the ONNX INT4 model and follow the CPU ONNX Runtime guide.

High-throughput serving: SGLang Omni

Audio8 TTS now includes an SGLang Omni adapter for production-oriented GPU serving. It is installed as an independent model plugin and does not overwrite SGLang Omni core files.

CapabilitySupport
Attention and batchingSGLang paged attention and dynamic batching
DualAR executionSlow AR serving with a fixed KV cache for the Fast AR codebook decoder
Voice cloningReference-audio encoding and waveform decoding
APIOpenAI-compatible /v1/audio/speech service

The released adapter is validated against a pinned SGLang Omni revision and supports both generation without a reference and zero-shot voice cloning. See the SGLang Omni deployment guide for tested versions, installation, server configuration, and API examples.

Evaluation

Audio8 TTS Preview is the smallest model in this comparison at just 0.6B parameters. Despite using only a fraction of the parameters of the other systems, it delivers results in the first tier of industry-leading SOTA TTS models on the benchmarks below. In particular, it achieves the best English WER and competitive Chinese CER on Seed-TTS, while remaining competitive across the CV3 multilingual evaluation.

Lower WER/CER is better; higher SIM is better. Seed-TTS similarity values are shown as percentages.

Seed-TTS

ModelParametersEN WER / SIMZH CER / SIMHard ZH CER / SIM
Audio8 TTS Preview0.6B1.506 / 63.20.950 / 73.111.510 / 68.7
Fish S2 Pro4.6B1.607 / 64.61.038 / 73.810.149 / 70.1
Higgs Audio v24.7B1.524 / 66.40.806 / 72.110.622 / 69.3
CosyVoice3-1.5B1.5B2.22 / 72.01.12 / 78.15.83 / 75.8
MOSS-TTS8.5B1.85 / 73.41.20 / 78.8-
VoxCPM22.3B1.84 / 75.30.97 / 79.58.13 / 75.3

CV3 multilingual error rate

ModelParameterszhenhard-zhhard-enjakodeesfritru
Audio8 TTS Preview0.6B3.2053.12810.5355.9977.2054.2233.4473.6418.7904.790-
Fish S2 Pro4.6B3.6003.49310.5887.3495.1394.1113.6052.9728.6004.2294.702
Higgs Audio v24.7B3.3783.40410.4245.7544.7424.2603.3002.9299.4253.5555.423
CosyVoice3-1.5B1.5B3.914.999.7710.557.575.696.434.4711.810.56.64
VoxCPM22.3B3.655.008.558.485.965.694.773.809.854.255.21

Parameter counts are calculated directly from the released weight tensors. MOSS-TTS contains 8,489,841,664 parameters. VoxCPM2's main model contains 2,290,004,544 parameters; the separate AudioVAE is not included in the parameter comparison.

Fish S2 Pro was reevaluated because its official evaluation uses its own normalizer. Higgs Audio v2 was evaluated locally because concrete values were unavailable. All other baseline values were collected from their official reports through the VoxCPM repository.

Different normalizers and evaluators make cross-project values reference comparisons rather than a strictly matched ranking. Evaluation coverage does not expand the Preview checkpoint's supported-language claim beyond the 11 languages listed above.

Limitations and Responsible Use

  • This is a Preview checkpoint with limited multilingual and dialect coverage.
  • Very long, noisy, or incorrectly transcribed reference clips can reduce stability and speaker similarity.
  • Generated speech can be misused for impersonation or misinformation. Obtain consent before cloning a voice and clearly disclose synthetic audio where appropriate.
  • Evaluate the model for accuracy, safety, and legal compliance before deployment.

License and Acknowledgements

The code and model weights are released under the Apache License 2.0. See the upstream NOTICE for attribution details.

We thank the Fish Audio team for publishing the DualAR architecture used in Fish Audio S2 Pro.