Kimi K3 on Fireworks: Frontier Intelligence You Can Own

Pricing to seamlessly scale from idea to enterprise

Start building in seconds, self-serve. Contact us for enterprise deployments with faster speeds, lower costs, and higher rate limits.

Serverless Inference

Pay per token, with high rate limits and postpaid billing. Get started with $1 in free credits. To view the current pricing for our most popular models across Standard, Priority, and Fast serverless tiers, visit our documentation.


Embeddings

Base model parameter count$ / 1M input tokens
up to 150M$0.008
150M - 350M$0.016
Qwen3 8B$0.1

Training Pricing

Serve fine-tuned models for the same price as base models.

Managed Training

Supervised and preference fine tuning is priced per 1M training tokens. Reinforcement fine tuning jobs are priced per GPU hour (billed per second), at the same price as Fireworks on-demand deployments (see on-demand pricing below).

Base ModelLoRA SFTLoRA DPOFull Param SFTFull Param DPO
Models up to 16B parameters$0.50$1.00$1.00$2.00
Models 16.1B - 80B$3.00$6.00$6.00$12.00
Models 80B - 300B (e.g. Qwen3-235B, gpt-oss-120B)$6.00$12.00$12.00$24.00
Models >300B (e.g. DeepSeek V3, Kimi K2)$10.00$20.00$20.00$40.00
  • SFT and DPO prices are shown in $ per 1M training tokens. Training tokens can be estimated with number of tokens in training dataset * number of epochs. Estimation should be multiplied by the average number conversation turns /2 for tuning with intermediate thinking traces.
  • Please note that when fine-tuning with reasoning traces, including the reasoning_content field for assistant turns will increase the total number of tuned tokens because multi-turn conversations are unrolled into user, assistant, and thinking traces. For further details, please refer to example 2 in the documentation about SFT fine tuning.
  • Fine-tuning with images (VLM supervised fine-tuning) is also billed per 1M tokens. See this FAQ on calculating image tokens.

Serverless Training API

Attach to a shared, always-on trainer pool for LoRA training on the launch models. There's no provisioning and no idle cost. You pay only for the tokens you prefill, sample, and train.

Base ModelContextPrefill / 1MCached Prefill / 1M Sample / 1MTrain / 1M
Qwen 3.5 9B65K$0.66$0.132$1.995$1.463
Qwen 3.6 27B65k$1.86$0.372$5.595$4.103
Kimi K3192k$10.87$2.17$27.11$32.55
  • Checkpoint storage for serverless models is included during private preview.
  • Other frontier models are coming soon to the Serverless Training API catalog.

Dedicated Training API

Dedicated Training API jobs are priced per GPU hour. Please see the On-Demand Pricing section below for details on Dedicated Training API Pricing.


On-Demand Pricing

Pay per GPU second, with no extra charges for start-up times

On demand deployments

GPU TypePrice ($) per hour
H100 80 GB GPU$7.00
H200 141 GB GPU$7.00
B200 180 GB GPU$10.00
B300 288 GB GPU$12.00
  • For estimates of per-token prices, see this blog. Results vary by use case, but we often observe improvements like ~250% higher throughput and 50% faster speed on Fireworks compared to open source inference engines.