Alpha
All entriesContact
Inference provider

Baseten

HQ
United States
Serves
9 open model families
Pricing
mixed
API
OpenAI-compatible

You substantially own inference here: an independent Cowork browser session confirmed, on two separate fetches, that Baseten's Terms & Conditions state Zero Data Retention by default, a contractual bar on using your content (including deployed model weights) to train its models, and a genuinely mutual confidentiality duty - a stronger standard-terms posture than most peers, and now independently verified rather than resting on a single automated read. A real, committed 99.9% SLA with service credits was also confirmed. It falls short of full ownership because a prior ISO 27001 claim was retracted (not found anywhere on independent verification), the Vanta Trust Center is fully JS-gated so PCI/FedRAMP/CSA STAR status and the named sub-processor list remain unverified, no published EU region exists for the managed service (EU residency means the enterprise Self-Hosted tier), and Async Inference has a 72-hour retention exception. If you need EU data residency or want certifications independently confirmed, use the Self-Hosted tier or ask for the underlying reports directly.

Do you really own it?
Substantial
none·limited·partial·substantial·full
Analytical input: AOI B · 73.6/100
The four ownership factors

Floor-weighted, not averaged. Nothing is weak and both use & modify and data control are strong, so the floor is high; transparency sits at moderate, which is what keeps it short of full.

1

Use and modify freelyCan you use it freely and leave without lock-in?

Strong

OpenAI/Anthropic-compatible APIs over a predominantly open-weight Model Library keep workloads portable, and the Terms & Conditions confirm the customer retains all IP in Customer Content including deployed model weights and outputs - the three-tier deployment ladder (shared/dedicated/self-hosted) gives real control over where and how the workload runs.

How this scores (AOI sub-dimensions)
Transparency & lock-in4/5how portable it is and how easily you can leaveAn OpenAI-compatible API (plus a beta Anthropic-compatible API) over a predominantly open-weight catalogue keeps workloads portable, and the Terms & Conditions give a documented 30-day post-termination export window plus an on-request deletion right.
Cost4/5how the pricing model compares and how predictable it isDetailed, public per-token rates for Model APIs and per-GPU-minute rates across a wide range of GPU classes (T4 through B200) are both published directly on Baseten's pricing page, with clear tier definitions (Basic/Pro/Enterprise).
2

TransparencyAre the binding terms published, legible and independently checkable?

Moderate

The Terms & Conditions and DPA are legible and were independently browser-read (Sec 11 mutual confidentiality confirmed on two separate fetches) - but a prior ISO 27001 claim was RETRACTED (not found anywhere, not even as a self-attestation) and the Vanta Trust Center that would carry PCI/FedRAMP/CSA STAR status is fully JS-gated and unreadable, so the sub-processor list's actual named contents also remain unverified.

How this scores
Not a scored AOI dimension. For a hosted provider, transparency is whether the binding terms are published, legible and were actually read - the read/unverified evidence below, not a certification. A strong rating here must trace to a retrieved binding document.
3

ReliabilityDoes it stay up and stay secure?

Strong

UPGRADED 2026-09-21: a real, committed SLA was independently browser-read (99.9% uptime, 40%-capped service credits, 24h claim window) - a prior draft held this at moderate because no SLA had been located. The public status page shows all systems operational with three brief, same/next-day-resolved incidents in the current window.

How this scores (AOI sub-dimensions)
Reliability4/5whether it stays up, with an SLA and status historyUPGRADED 2026-09-21: an independent Cowork browser session found and directly read a real, committed SLA (baseten.co/service-level-agreement, Sec 2.0) - 99.9% uptime for both Dedicated Inference and Model APIs, with a service-credit remedy (capped at 40% of monthly fees, Sec 4.2) and a 24-hour claim window (Sec 4.3).
Security4/5the controls protecting your traffic and dataCORRECTED 2026-09-21 (description, not score): independently browser-confirmed isolation is logical namespace separation for shared use (access controls, unique customer identifiers, namespace isolation) plus single-tenant dedicated Kubernetes namespaces with Calico/Cilium network policies ("GPUs never shared across users") and self-hosted deployment inside the customer's own VPC - not an explicit "three-tier" model as an earlier draft described it, though the practical effect is similar.
Compliance4/5which independent certifications and attestations it holdsCORRECTED 2026-09-21: SOC 2 Type II is confirmed with a named audit firm (Sensiba San Filippo LLP, clean opinion across Security/Availability/Processing Integrity/Confidentiality/Privacy criteria); HIPAA and GDPR are independently confirmed on the Security Practices and DPA pages.
4

Doesn't extract your dataDo the binding terms keep your data and IP yours?

Strong

The Terms & Conditions and security-practices page (independently browser-read, Sec 11 confirmed on two fetches) state Zero Data Retention by default for Model APIs and Dedicated Inference, a contractual never-train clause, and a MUTUAL confidentiality duty - among the strongest standard-terms data-control postures reviewed, now independently verified rather than resting on a single automated read. Caveat: Async Inference queues inputs for up to 72 hours, and no published EU region exists for the managed service.

How this scores (AOI sub-dimensions)
Data governance4/5retention, training-on-inputs and data ownershipTerms & Conditions and the security-practices/DPA pages together give a strong, binding posture: Zero Data Retention by default for Model APIs and Dedicated Inference, a contractual never-train clause covering Customer Content (including deployed model weights and outputs), customer ownership of Customer Content, and a MUTUAL confidentiality duty (same care as own information, three-year survival).
Residency2/5where your data is processed and storedA "regional environments" feature exists, routing a deployment exclusively to a designated geographic region, but no published list of available regions - including whether an EU region is among them - was found; the docs direct customers to contact support.

How the AOI score is computed

The seven dimensions above, each scored 0 to 5, weighted and summed to the 0 to 100 headline. The score is the analytical input behind the ownership verdict, not the verdict itself.

DimensionScoreWeightPoints
Data governance4/50.2419.2
Compliance4/50.1814.4
Residency2/50.166.4
Security4/50.1411.2
Reliability4/50.129.6
Transparency & lock-in4/50.108.0
Cost4/50.064.8
HeadlineB · 73.6/100
Dossier coverageAssess 81%Implement 50%Use 50%Support 43%How complete our four-domain documentation is, a measure of our coverage, not of the model. Each domain links to its page.

Sources

Every rating traces to a primary document. Read means the text was verified; unverified means it is known to exist but has not yet been read.

DocumentWhat it grounds
Terms of serviceread2026-09-20
Baseten Terms & Conditions (read directly): Sec on Customer Content ownership - 'Customer reserves all rights, title, and interest in and to Customer Content,' where Customer Content includes Customer Models (deployed weights) and Model Outputs; a training-prohibition clause bars Baseten from using Customer Content to train, fine-tune, or otherwise develop ML/AI models, limiting its license to providing/maintaining the Services and preventing/addressing technical problems, plus anonymized/de-identified telemetry; mutual confidentiality clause requiring 'the same degree of care it uses for its own (but no less than reasonable care),' surviving three years post-termination, with prior-notice compelled disclosure; a documented exit right - Customer Content exportable for 30 days after Term end, deletable on written request.
Securityread2026-09-20
Baseten security-practices page (read directly): Zero Data Retention for Model APIs and Dedicated Inference ('will not store, retain, or otherwise make a persistent copy of model inputs or outputs'); Async Inference queues inputs up to 72 hours before deletion, outputs never stored; three-tier tenant isolation (shared logical separation; dedicated single-tenant Kubernetes namespaces with Calico/Cilium network policies, 'never shares GPUs across users'; self-hosted in customer VPC); TLS 1.2+ in transit, AES-256 at rest with periodic key rotation; least-privilege access, enforced MFA, quarterly access reviews; periodic third-party penetration testing; KV cache never persisted to disk.
Data Processing Addendumread2026-09-20
Baseten DPA (read directly): 'Baseten shall not log, record, or save Customer Personal Data contained in model inputs or outputs to persistent storage after real-time processing'; Baseten as Processor, Customer as Controller; SCCs incorporated, governing law/jurisdiction Ireland (Clause 17/18); UK Addendum (ICO template B.1.0) for UK transfers; a sub-processor list is maintained with at least fifteen (15) days' prior notice before engaging a new sub-processor.
Vendor announcementread2026-09-20
Baseten blog post confirming SOC 2 Type II certification, audited by Sensiba San Filippo LLP, covering Security/Availability/Processing Integrity/Confidentiality/Privacy criteria with a clean opinion across 70+ sub-criteria; the report itself is gated under NDA.
Documentationread2026-09-20
Baseten regional-environments docs (read directly): routes inference traffic for a deployment 'exclusively to workload planes within a designated geographic region'; requires initial configuration by Baseten - 'Contact support to confirm availability for the region you need.' No published list of specific region names was found.
Documentationread2026-09-20
Baseten Self-Hosted product page (read directly): the full workload plane runs inside the customer's own VPC/cloud, with inference input/output 'never touching Baseten's premises'; marketed toward the Enterprise pricing tier.
Slaread2026-09-20
UPGRADED 2026-09-21: Baseten's public status page (status.baseten.co), independently browser-read: 'All Systems Operational' across 6 components (Dedicated Inference, Model APIs, Training, Model Management API, Web Application, Homepage and Docs); 90-day history, no numeric % shown on the face.
Slaread2026-09-21
Baseten Service Level Agreement (independently browser-read 2026-09-21, verbatim): Sec 2.0 - 'ninety-nine point nine percent (99.9%)' committed uptime, applying to both Dedicated Inference and Model APIs on Baseten-managed infrastructure.
Vendor announcementread2026-09-20
Baseten pricing page (read directly): per-token Model API rates (e.g.
Documentationread2026-09-20
Baseten Model Library page (read directly): current curated catalogue includes DeepSeek V4.1 Flash, GLM-5/5.3/5.3 Fast, Llama 3.3 70B Instruct, Kimi K3, Qwen3.5 35B-A3B + TTS + Reranker/Embedding + Image, Whisper Large V3 variants, Flux.2 [dev], EmbeddingGemma, Nomic Embed Code, BGE Embedding ICL, Inkling.
Documentationread2026-09-20
Baseten platform overview docs (read directly): three pathways - hosted Model APIs supporting the OpenAI Chat Completions API and a beta Anthropic Messages API, deploying an open-source/fine-tuned/custom model on dedicated GPUs, and Training Jobs/Loops.