Models served
Fireworks' model catalogue (fireworks.ai/models, read directly) is broad and fast-moving:
- Kimi K3 / K2.7 (Code) / K2.6
- DeepSeek V4.1 Flash / V4 Pro / V4 Flash (incl. Vision Exp)
- GLM-5.3 (Flash/Fast/US variants) / GLM-5.2 / GLM-4.5V (vision)
- Qwen3.8 Max/Flash, Qwen3 VL, Qwen3 Embedding/Reranker
- Llama 3.3 70B, 3.2 (incl. Vision), 3.1 family
- Mistral Large 3 / Ministral 3 / Nemo
- NVIDIA Nemotron 3 Ultra / 3.5 Lightning
- gpt-oss 120B
- Embeddings (BGE-M3, Voyage 4) and image (FLUX.1 schnell/dev + ControlNet)
Recheck this list periodically - the catalogue moves generations quickly, as it did between the
prior and current passes on peer providers in this library.
Inference features
Endpoints are OpenAI-compatible chat/completions, plus an OpenAI-compatible Responses API
with MCP support. Per-model tool-calling and structured-output support follows each served
model's own capabilities and was not individually verified on Fireworks' endpoints - confirm per
model.
Fine-tuning / custom models
Fireworks documents on-demand dedicated deployment of custom and fine-tuned models, but the
specific fine-tuning workflow - supported methods, dataset formats, base models - was not read in
full this pass. Confirm the current scope in the Fireworks docs before planning a custom-model
workflow.
Context & token limits
Fireworks does not publish a single per-model context-limit table in the sources read; limits
track each served model's native window and should be confirmed on the model's own catalogue
page.
The ownership factor this domain covers, drawn from the one entry record.
3
ReliabilityDoes it stay up and stay secure?
ModerateA public status page (independently browser-confirmed) shows a strong recent uptime record (~99.67-100% across 28 tracked endpoints), but no contractual SLA or service-credit clause was found, and no long multi-year track record was independently verified.
How this scores (AOI sub-dimensions)
Reliability3/5whether it stays up, with an SLA and status historyA public status page (status.fireworks.ai) shows an operational track record with per-endpoint 90-day uptime around 99.67-100% and documented incident communication.
Security3/5the controls protecting your traffic and dataDedicated workloads run in "logically isolated environments, preventing cross-customer access or data leakage" per the data-security docs (independently browser-confirmed verbatim), with TLS 1.2+ in transit and AES-256 at rest, least-privilege access controls, and customer-managed keys (CMEK).
Compliance4/5which independent certifications and attestations it holdsFireworks' own docs certifications FAQ and blog announce SOC 2 Type II (2023-10-27) plus HIPAA compliance, and a separate blog announces triple ISO certification (27001/27701/42001, 2025-11-19) - both independently browser-read; a trust centre exists at trust.fireworks.ai.
The same evidence records as the entry sheet. Read means the text was verified; unverified means it is known to exist but not yet read.
Terms of serviceunverified2026-09-20
UNCONFIRMED as of 2026-09-20.
Privacy Policyread2026-09-20
Fireworks Privacy Policy (independently browser-read 2026-09-20, lastmod 2026-08-11): 'We do not use your prompts, training data, or API inputs to train or improve our AI models without your explicit opt-in.' 'We do not log or store prompt or generation data for any open models without explicit user opt-in.' This is Privacy Policy language, not a section-numbered ToS clause, but it independently corroborates the substance of no-training-by-default via a readable document.
Documentationread2026-09-20
Fireworks data-handling docs (independently browser-read 2026-09-20): 'The Response API operates under a different retention model when store=True (the default setting).' 'Stored conversation data automatically deletes after 30 days.' 'Users can prevent storage by setting store=False in API requests.' 'The DELETE API endpoint enables immediate removal of specific records by providing the response_id.' 'The Response API retention policy only applies to conversation data when using the Response API endpoints.
Data Processing Addendumunverified2026-09-20
UNCONFIRMED as of 2026-09-20.
Securityread2026-09-20
Fireworks data-security docs (independently browser-read 2026-09-20, verbatim): 'Data is encrypted in transit (TLS 1.2+) and at rest (AES-256).' 'Dedicated workloads run in logically isolated environments, preventing cross-customer access or data leakage.' 'Fine-grained access controls are enforced across all Fireworks environments, following the principle of least privilege.' 'Regular penetration testing validates controls.' No bug bounty programme mentioned.
Vendor announcementread2026-09-20
Fireworks blog post announcing SOC 2 Type II certification and HIPAA compliance.
Vendor announcementread2026-09-20
Fireworks blog post announcing triple ISO certification: ISO 27001, ISO 27701, and ISO 42001.
Securityunverified2026-09-20
Fireworks' Trust Center portal exists at trust.fireworks.ai; it is a JS-rendered SafeBase page whose certificate/report contents were not independently retrieved this pass.
Documentationread2026-09-20
Fireworks serverless pricing docs (read directly): per-token pricing (input / cached input at a discount / output), batch inference at 50% of standard price, generic (non-featured) models priced by parameter-size band, and a 'US-only serverless' residency tier at a 1.5x premium (effective 2026-09-01).
Documentationread2026-09-20
Fireworks on-demand/dedicated-deployment docs (read directly): self-serve deployment via firectl/REST/SDKs/console, billed per GPU-second; regions are GLOBAL by default, with US/Europe/APAC/single-region pinning available but requiring a sales-granted quota.
Vendor announcementread2026-09-20
Fireworks 'Virtual Cloud' (BYOC) blog post (read directly): the Fireworks inference engine runs inside the customer's own VPC so 'data never leaves your secure environment'; GA announced 2025-06-16; access is enterprise/'contact us', not self-serve.
Slaread2026-09-20
Fireworks public status page (read directly): fully operational at check time, with per-endpoint 90-day uptime figures (~99.67-100% range) across roughly 28 tracked serverless model endpoints; no SLA/credit terms are published there.
Documentationread2026-09-20
Fireworks model catalogue page (read directly): current top models include Kimi K3/K2.7/K2.6, DeepSeek V4.1 Flash/V4 Pro/V4 Flash, GLM-5.3/5.2/4.5V, Qwen3.8 Max/Flash + VL + embeddings/reranker, Llama 3.x, Mistral Large 3/Ministral 3/Nemo, NVIDIA Nemotron 3 Ultra/3.5 Lightning, gpt-oss-120B, plus BGE-M3/Voyage embeddings and FLUX image models.
Documentationread2026-09-20
Fireworks OpenAI-compatibility docs (independently browser-read 2026-09-20): base URL 'https://api.fireworks.ai/inference/v1', confirmed OpenAI-compatible Chat Completions and Completions support.