Alpha
Contact
Use · Baseten

Is it good enough?

Ownership levelSubstantialnone·limited·partial·substantial·fullAnalytical input B ยท 73.6/100

This page is a projection of the one entry record, the Reliability factor that Use covers. The full verdict is set by all four factors together, floor-weighted so the weakest caps the whole.

Which domain expands which factor
  • AssessUse & modify + Transparency
  • ImplementData control + Reliability
  • UseReliability
  • SupportTransparency

Models served

Baseten's curated Model Library (baseten.co/library, read directly) currently features:

  • DeepSeek V4.1 Flash (552B MoE, multimodal, 1M-token context)
  • GLM-5 / GLM-5.3 / GLM-5.3 Fast
  • Kimi K3
  • Llama 3.3 70B Instruct
  • Qwen3.5 35B-A3B, Qwen3 TTS, Qwen3 8B Reranker/Embedding, Qwen Image
  • Whisper Large V3 (transcription)
  • Flux.2 [dev] (image generation)
  • Embeddings: EmbeddingGemma, Nomic Embed Code, BGE Embedding ICL

Beyond the curated library, Baseten is equally a bring-your-own-model deploy platform: any open-source, fine-tuned, or fully custom model can be deployed on dedicated GPUs.

Inference features

Hosted Model APIs support the OpenAI Chat Completions API and a beta Anthropic Messages API. Per-model tool-calling and structured-output support follows each served model's own capabilities and was not individually verified on Baseten's endpoints - confirm per model.

Fine-tuning / custom models

Training Jobs/Loops are documented as a platform pathway alongside Model APIs and custom dedicated-GPU deploys (docs.baseten.co/overview), but the specific fine-tuning workflow was not read in full this pass - confirm supported methods and base models before planning a custom-model workflow.

Context & token limits

Baseten does not publish a single per-model context-limit table in the sources read; limits track each served model's native window (DeepSeek V4.1 Flash's 1M-token context is noted on its own model page) and should be confirmed per model.

How this scores

The ownership factor this domain covers, drawn from the one entry record.

3

ReliabilityDoes it stay up and stay secure?

Strong

UPGRADED 2026-09-21: a real, committed SLA was independently browser-read (99.9% uptime, 40%-capped service credits, 24h claim window) - a prior draft held this at moderate because no SLA had been located. The public status page shows all systems operational with three brief, same/next-day-resolved incidents in the current window.

How this scores (AOI sub-dimensions)
Reliability4/5whether it stays up, with an SLA and status historyUPGRADED 2026-09-21: an independent Cowork browser session found and directly read a real, committed SLA (baseten.co/service-level-agreement, Sec 2.0) - 99.9% uptime for both Dedicated Inference and Model APIs, with a service-credit remedy (capped at 40% of monthly fees, Sec 4.2) and a 24-hour claim window (Sec 4.3).
Security4/5the controls protecting your traffic and dataCORRECTED 2026-09-21 (description, not score): independently browser-confirmed isolation is logical namespace separation for shared use (access controls, unique customer identifiers, namespace isolation) plus single-tenant dedicated Kubernetes namespaces with Calico/Cilium network policies ("GPUs never shared across users") and self-hosted deployment inside the customer's own VPC - not an explicit "three-tier" model as an earlier draft described it, though the practical effect is similar.
Compliance4/5which independent certifications and attestations it holdsCORRECTED 2026-09-21: SOC 2 Type II is confirmed with a named audit firm (Sensiba San Filippo LLP, clean opinion across Security/Availability/Processing Integrity/Confidentiality/Privacy criteria); HIPAA and GDPR are independently confirmed on the Security Practices and DPA pages.
What this means for adoptionYou substantially own inference here: an independent Cowork browser session confirmed, on two separate fetches, that Baseten's Terms & Conditions state Zero Data Retention by default, a contractual bar on using your content (including deployed model weights) to train its models, and a genuinely mutual confidentiality duty - a stronger standard-terms posture than most peers, and now independently verified rather than resting on a single automated read. A real, committed 99.9% SLA with service credits was also confirmed. It falls short of full ownership because a prior ISO 27001 claim was retracted (not found anywhere on independent verification), the Vanta Trust Center is fully JS-gated so PCI/FedRAMP/CSA STAR status and the named sub-processor list remain unverified, no published EU region exists for the managed service (EU residency means the enterprise Self-Hosted tier), and Async Inference has a 72-hour retention exception. If you need EU data residency or want certifications independently confirmed, use the Self-Hosted tier or ask for the underlying reports directly.

Sources

The same evidence records as the entry sheet. Read means the text was verified; unverified means it is known to exist but not yet read.

Terms of serviceread2026-09-20
Baseten Terms & Conditions (read directly): Sec on Customer Content ownership - 'Customer reserves all rights, title, and interest in and to Customer Content,' where Customer Content includes Customer Models (deployed weights) and Model Outputs; a training-prohibition clause bars Baseten from using Customer Content to train, fine-tune, or otherwise develop ML/AI models, limiting its license to providing/maintaining the Services and preventing/addressing technical problems, plus anonymized/de-identified telemetry; mutual confidentiality clause requiring 'the same degree of care it uses for its own (but no less than reasonable care),' surviving three years post-termination, with prior-notice compelled disclosure; a documented exit right - Customer Content exportable for 30 days after Term end, deletable on written request.
Securityread2026-09-20
Baseten security-practices page (read directly): Zero Data Retention for Model APIs and Dedicated Inference ('will not store, retain, or otherwise make a persistent copy of model inputs or outputs'); Async Inference queues inputs up to 72 hours before deletion, outputs never stored; three-tier tenant isolation (shared logical separation; dedicated single-tenant Kubernetes namespaces with Calico/Cilium network policies, 'never shares GPUs across users'; self-hosted in customer VPC); TLS 1.2+ in transit, AES-256 at rest with periodic key rotation; least-privilege access, enforced MFA, quarterly access reviews; periodic third-party penetration testing; KV cache never persisted to disk.
Data Processing Addendumread2026-09-20
Baseten DPA (read directly): 'Baseten shall not log, record, or save Customer Personal Data contained in model inputs or outputs to persistent storage after real-time processing'; Baseten as Processor, Customer as Controller; SCCs incorporated, governing law/jurisdiction Ireland (Clause 17/18); UK Addendum (ICO template B.1.0) for UK transfers; a sub-processor list is maintained with at least fifteen (15) days' prior notice before engaging a new sub-processor.
Vendor announcementread2026-09-20
Baseten blog post confirming SOC 2 Type II certification, audited by Sensiba San Filippo LLP, covering Security/Availability/Processing Integrity/Confidentiality/Privacy criteria with a clean opinion across 70+ sub-criteria; the report itself is gated under NDA.
Documentationread2026-09-20
Baseten regional-environments docs (read directly): routes inference traffic for a deployment 'exclusively to workload planes within a designated geographic region'; requires initial configuration by Baseten - 'Contact support to confirm availability for the region you need.' No published list of specific region names was found.
Documentationread2026-09-20
Baseten Self-Hosted product page (read directly): the full workload plane runs inside the customer's own VPC/cloud, with inference input/output 'never touching Baseten's premises'; marketed toward the Enterprise pricing tier.
Slaread2026-09-20
UPGRADED 2026-09-21: Baseten's public status page (status.baseten.co), independently browser-read: 'All Systems Operational' across 6 components (Dedicated Inference, Model APIs, Training, Model Management API, Web Application, Homepage and Docs); 90-day history, no numeric % shown on the face.
Slaread2026-09-21
Baseten Service Level Agreement (independently browser-read 2026-09-21, verbatim): Sec 2.0 - 'ninety-nine point nine percent (99.9%)' committed uptime, applying to both Dedicated Inference and Model APIs on Baseten-managed infrastructure.
Vendor announcementread2026-09-20
Baseten pricing page (read directly): per-token Model API rates (e.g.
Documentationread2026-09-20
Baseten Model Library page (read directly): current curated catalogue includes DeepSeek V4.1 Flash, GLM-5/5.3/5.3 Fast, Llama 3.3 70B Instruct, Kimi K3, Qwen3.5 35B-A3B + TTS + Reranker/Embedding + Image, Whisper Large V3 variants, Flux.2 [dev], EmbeddingGemma, Nomic Embed Code, BGE Embedding ICL, Inkling.
Documentationread2026-09-20
Baseten platform overview docs (read directly): three pathways - hosted Model APIs supporting the OpenAI Chat Completions API and a beta Anthropic Messages API, deploying an open-source/fine-tuned/custom model on dedicated GPUs, and Training Jobs/Loops.