Models served
Nebius' live catalogue (tokenfactory.nebius.com/models/catalog, browsed directly) runs to roughly
90 model variants, including:
- DeepSeek-V4-Pro, DeepSeek-V4-Flash, DeepSeek-R1-0528, DeepSeek-V3-0324
- Kimi-K2-Instruct, Kimi-K2-Thinking, Kimi-K3
- Qwen3-235B-A22B, Qwen3-Coder-480B-A35B
- Llama-3.3-70B-Instruct, Llama-3.1-8B-Instruct, Llama-Guard-3-8B
- gpt-oss-120b / gpt-oss-20b
- GLM-4.5 / GLM-5.1 / GLM-5.2 / GLM-5.3
- NVIDIA Nemotron-3-Ultra-550B-A55B, Nemotron-3-Nano-30B-A3B
- MiniMax-M2 / M3
Correction (2026-09-21): an earlier check found several of the models above -
Kimi-K2-Instruct, Llama-3.3-70B-Instruct, and GLM-4.5 - showing "Public endpoint: Not
available." An independent re-check found this NOT to mean dedicated-only access: those
specific named versions have simply been superseded by newer catalogue entries (e.g.
Kimi-K2.7-Code, GLM-5.x) - ordinary catalogue churn. Open models are confirmed publicly served
per-token (DeepSeek-V4-Pro, gpt-oss-120b, and Llama-3.3-70B-Instruct at $0.13/$0.40 per 1M
tokens, all independently confirmed). The catalogue moves quickly - confirm current
public-endpoint availability on the model's own page.
Inference features
Endpoints are OpenAI-compatible chat/completions. Per-model tool-calling and
structured-output support follows each served model's own capabilities and was not individually
verified on Nebius' endpoints - confirm per model.
Fine-tuning / custom models
A fine-tuning feature is referenced in the Legal Quick Guide, which states fine-tuning artifacts
are stored exclusively in EU data centres - but the specific workflow (supported methods,
base models) was not read in full this pass.
Context & token limits
Nebius does not publish a single per-model context-limit table in the sources read; limits track
each served model's native window and should be confirmed on the model's own catalogue page.
The ownership factor this domain covers, drawn from the one entry record.
3
ReliabilityDoes it stay up and stay secure?
ModerateCORRECTED 2026-09-21 (downward): a prior claim of a real contractual SLA with service credits is RETRACTED - an independent browser session found no committed inference SLA exists at all; the master SLA page delegates to seven per-service sub-pages, none of which is inference. What remains: a public status page with a dedicated Token Factory component, shown operational, with a documented but informal incident record.
How this scores (AOI sub-dimensions)
Reliability3/5whether it stays up, with an SLA and status historyCORRECTED 2026-09-21: an independent Cowork browser session found NO committed inference SLA at all - a prior draft claimed a real contractual SLA existed, which overstated it.
Security3/5the controls protecting your traffic and dataEncryption at rest and in transit and need-to-know access management are referenced in the Legal Quick Guide (read directly), dedicated endpoints are marketed as single-tenant/isolated, and SOC 2 Type II implies independently tested operational controls.
Compliance4/5which independent certifications and attestations it holdsSTRENGTHENED 2026-09-21: independently confirmed verbatim - SOC 2 Type II is audited by a NAMED firm, "Deloitte", covering HIPAA; the ISO 27001 scope statement explicitly names "AI Cloud, AI Studio and TractoAI" (AI Studio = Token Factory, confirming this product is actually in scope, not just the company generally).
The same evidence records as the entry sheet. Read means the text was verified; unverified means it is known to exist but not yet read.
Terms of serviceread2026-09-20
Nebius Token Factory Terms of Service (read directly), Sec 7: Nebius collects and stores Inputs/Outputs by default and grants itself a licence to 'access, use, host, cache, store, copy, and modify Inputs and Outputs' to provide the Service AND to train its own smaller models used exclusively for 'Speculative Decoding'; opt-out available via onboarding form or emailing tokenfactory-support@nebius.com; separately reserves the right to 'remove, screen, or delete any of Your Inputs and Outputs at any time, for any reason, and without notice.' Sec 10(b)/(c): customer 'holds exclusive ownership of all rights, titles, and interests...
Documentationread2026-09-20
Nebius Legal Quick Guide (read directly): states 'We do not use your content to train, fine-tune or improve any AI models - ours or third parties'' - in tension with the binding Terms of Service Sec 7 (see ev-tos).
Terms of serviceread2026-09-20
Nebius Master Services Agreement (read directly), Sec 7.11: 'Nebius may use information about how the Customer use and interacts with the Services for the purpose of improvement of the Services...
Data Processing Addendumread2026-09-20
Nebius Data Processing Addendum (read directly): imposes a processor-side confidentiality duty specifically over 'Customer Personal Data' ('any person that it authorizes to process Customer Personal Data...
Subprocessorsread2026-09-20
CORRECTED 2026-09-21 (count): Nebius Token Factory sub-processor list, independently browser-read, effective 2026-09-15: approximately 21 named entities enumerated (the page's own summary line said 18; a prior draft said ~24 - an exact recount is recommended) across Nebius Group entities (Nebius Inc.
Securityread2026-09-20
STRENGTHENED 2026-09-21: Nebius Trust Center and the linked SOC 2 blog post, both independently browser-read, verbatim: SOC 2 Type II auditor is NAMED - 'Deloitte, an accredited third-party firm that evaluates the design and operational effectiveness of our measures to protect customer data' - and 'SOC 2 Type II also includes a section that confirms compliance with...
Documentationread2026-09-20
Nebius regions documentation (read directly): public regions eu-north1 (Finland), eu-west1 (France), me-west1 (Israel), uk-south1/uk-south2 (UK), us-central1 (US); private/ negotiated regions eu-north2 (Iceland), eu-south1 (Madrid, Spain), eu-west2 (France), us-north1 (US).
Slaread2026-09-20
CORRECTED 2026-09-21: Nebius master SLA page, independently browser-read, verbatim: 'The list of Services which provides Service Levels and links for Service Levels for specific Service are available at: https://docs.nebius.com/legal/sla-levels' and 'Service Level and amount of Compensation is determined for each Service separately.' No number in the master doc; no mention of inference/Token Factory/AI Studio.
Slaread2026-09-20
UPGRADED 2026-09-21: Nebius public status page (status.nebius.com), independently browser-read: a dedicated 'Token Factory' component shown Operational (no separate 'Inference'/'AI Studio' component); regions covered EU-NORTH1, EU-NORTH2, EU-WEST1, EU-WEST2, UK-SOUTH1, US-CENTRAL1, ME-WEST1; 'Uptime over the past 90 days' shown with no numeric % on the main view.
Documentationread2026-09-20
CORRECTED 2026-09-21: Nebius Token Factory live model catalogue, re-checked independently (tokenfactory.nebius.com/models/catalog): roughly 90 model variants.
Documentationread2026-09-20
Nebius dedicated-endpoint billing-policy docs (read directly): billed per running replica on a pay-as-you-go basis, adjusting dynamically with autoscaling; 'charges may vary depending on your custom contract or work order' - the underlying per-replica/GPU-hour rate itself is not disclosed on this page.
Documentationread2026-09-20
Nebius Token Factory product page (read directly): 'a simple, OpenAI-compatible API' over an open-weight model catalogue, with both shared/public per-token endpoints and dedicated single-tenant endpoints.