All entriesContact
Model

OpenAI gpt-oss

Publisher
OpenAI (US)
Family
gpt-oss-120b / gpt-oss-20b
Openness
open_weights
Licence
Apache License 2.0
Context
128K

gpt-oss is OpenAI's first open-weight release since GPT-2 (August 2025): two mixture-of-experts reasoning models - gpt-oss-120b (~117B total / 5.1B active, 128 experts) and gpt-oss-20b (~21B total / 3.6B active, 32 experts) - both with a 128K context window, configurable low/medium/high reasoning effort, native agentic tool use (function calling, browsing, Python) and structured outputs, shipped under a clean OSI-approved Apache-2.0 license with no click-through gate.

Do you really own it?
Substantial
none·limited·partial·substantial·full
Analytical input: AOI B · 78.8/100

Floor-weighted, not averaged. The weakest factor caps the level, because ownership is a conjunction. No factor is weak and both use-and-modify and data-control are strong, so the level is substantial; a moderate elsewhere keeps it short of full.

You substantially own gpt-oss: ungated Apache-2.0 weights you run, modify and self-host, governed by a one-sentence usage policy (obey applicable law) that is about as light as an open-weight AUP gets. The safety posture is a genuine differentiator - a published worst-case malicious-fine-tuning study whose methodology METR independently reviewed, with OpenAI at least partially addressing each of METR's six high-urgency items. What holds it below full is transparency: the training data and pre-training code are closed, so you adapt rather than reproduce. Adopt it for general open-weight deployment, building to the harmony response format it requires and adding standard input/output guardrails for customer-facing use.

1

Use and modify freelyCan you run, modify and adapt it with no gate and no field-of-use trap?

Strong

A clean, OSI-approved Apache-2.0 grant - own, run, modify and redistribute freely with no access gate and no field-of-use limit - and the only usage constraint is a single-sentence policy asking you to obey applicable law, an unusually light AUP for an open-weight release.

How this scores (AOI sub-dimensions)
Openness3/5how much is released - weights, data, code, licence - and how freelyOpen-weights tier: ungated, downloadable Apache-2.0 weights with strong public documentation (model card + arXiv report), which is materially more open than a gated community license.
Legal4/5how permissive and clean the licence is for real commercial useA pristine, OSI-approved Apache-2.0 license - ungated, no acceptable-use policy, no MAU threshold, unconditional commercial use - well above a restrictive community license, and the model is under the systemic-risk FLOP threshold.
2

TransparencyDo you know what it is: weights, training, behaviour, and legible terms?

Moderate

Weights are open and inspectable and OpenAI published a safety report whose worst-case malicious-fine-tuning methodology was independently reviewed by METR, but the training data and pre-training code are closed, so you adapt rather than reproduce and cannot see how it was made.

How this scores (AOI sub-dimensions)
Provenance4/5how well we can trace and verify what went into the modelVerified openai org on Hugging Face, safetensors-only distribution with checksums, a clear canonical source and no malicious-checkpoint incident on the canonical org (checklist ~6/8).
Governance4/5how accountable and well-documented the publisher isReputable, legally accountable US publisher with a detailed model card / arXiv report, a public reference repository, a documented Preparedness-Framework process and coordinated external safety testing.
3

Doesn't fail youIs it reliable and good enough for the job?

Strong

Strong in class (headline 78.8) with 128K context and native tool use, backed by a first-class day-0 serving ecosystem across every major runtime and hosted endpoint, plus an independently-reviewed worst-case safety study.

How this scores (AOI sub-dimensions)
Performance4/5how capable it is relative to its classStrong reasoning models: gpt-oss-120b is reported near OpenAI o4-mini and gpt-oss-20b near o3-mini on core reasoning benchmarks, and they rank well among open-weight models on public leaderboards.
Operational5/5how practical it is to run, serve and maintain in productionFirst-class ecosystem support from day 0: safetensors on the verified openai org shipped in native MXFP4, with support across vLLM, llama.cpp, Ollama, Hugging Face transformers and MLX, plus hosted endpoints on major providers.
Safety4/5whether misuse risks are evaluated and guardrails are providedSafety-tuned reasoning releases with a published safety evaluation, Preparedness-Framework capability testing, and a worst-case malicious-fine-tuning study whose methodology was independently reviewed by external experts (METR, SecureBio, Daniel Kang).
4

Doesn't extract your dataDoes running it keep your knowledge and data yours?

Strong

Self-hosted, the ungated Apache weights run entirely on your own infrastructure with no clawback or telemetry - your data stays yours.

How this scores
Not a scored AOI dimension. For a self-hosted model, data-control is a structural property of running the weights yourself, strong by default unless the model phones home or the licence claws back rights. For a hosted API this factor is the retention + train-on-inputs + residency read, scored from the binding terms.

How the AOI score is computed

The seven dimensions above, each scored 0 to 5, weighted and summed to the 0 to 100 headline. The score is the analytical input behind the ownership verdict, not the verdict itself.

DimensionScoreWeightPoints
Openness3/50.1810.8
Provenance4/50.1612.8
Legal4/50.1612.8
Safety4/50.1612.8
Performance4/50.1411.2
Operational5/50.1212.0
Governance4/50.086.4
HeadlineB · 78.8/100
Dossier coverageAssess 97%Implement 100%Use 94%Support 50%How complete our four-domain documentation is, a measure of our coverage, not of the model. Each domain links to its page.

Sources

Every rating traces to a primary document. Read means the text was verified; unverified means it is known to exist but has not yet been read.

DocumentWhat it grounds
Vendor announcementunverified2026-07-25
OpenAI introduces gpt-oss-120b and gpt-oss-20b as open-weight reasoning models under Apache-2.0 with configurable reasoning effort, native agentic tool use and structured outputs, running on 80GB / 16GB hardware.
Technical_reportread2026-07-25
gpt-oss-120b & gpt-oss-20b Model Card (arXiv 2508.10925), read verbatim: two open-weight reasoning models on an efficient mixture-of-expert transformer architecture, trained with large-scale distillation and reinforcement learning, releasing all model weights, inference code, tools and tokenizers under an Apache-2.0 license; documents the MoE architecture (~117B/5.1B and ~21B/3.6B, 128/32 experts), 128K context, MXFP4, harmony format and the Preparedness-Framework safety and malicious-fine-tuning evaluations.
Model cardunverified2026-07-25
gpt-oss weights are hosted on the verified openai organisation on Hugging Face as safetensors under Apache-2.0 with a 128K context window.
Licenceread2026-07-25
gpt-oss LICENSE, read verbatim: a standard Apache License 2.0 (January 2004) boilerplate with the copyright line left as a template placeholder and no bespoke restrictions in the file itself - unconditional commercial use; any use guidance lives separately in the one-sentence USAGE_POLICY, not the licence.
Terms of serviceread2026-07-25
gpt-oss USAGE_POLICY, read verbatim and confirmed complete: the entire policy is a single sentence - "We aim for our tools to be used safely, responsibly, and democratically, while maximising your control over how you use them.
Third-party analysisunverified2026-07-25
Hugging Face's launch guide documents day-0 gpt-oss support across transformers, vLLM, Ollama and llama.cpp, the native MXFP4 quantization, and the harmony format.
Documentationunverified2026-07-25
gpt-oss is trained for OpenAI's harmony response format (role hierarchy and separate reasoning/tool/final channels), rendered via the openai/harmony library.
Securityread2026-07-25
OpenAI gpt-oss safety report PDF, read verbatim: adversarial malicious-fine-tuning on bio/cyber data to simulate a worst case (with external reviewers METR, SecureBio, Daniel Kang); the default model "does not reach High capability" in any of the three Tracked Categories, and OpenAI's Safety Advisory Group concluded that even robust fine-tuning "did not reach High capability in Biological and Chemical Risk or Cyber risk."
Third-party analysisread2026-07-25
METR third-party methodology review of OpenAI's gpt-oss malicious-fine-tuning safety work, read verbatim: "OpenAI at least partially addressed each of our 6 high-urgency items," while flagging that no published High-capability criteria exist; conducted alongside SecureBio and Daniel Kang.
Third-party analysisunverified2026-07-25
gpt-oss ranks strongly among open-weight models on independent public leaderboards.
Documentationunverified2026-07-25
The openai/gpt-oss reference repository provides implementations, serving guidance and fine-tuning references for the open weights.
Third-party analysisunverified2026-07-25
gpt-oss is available in the Ollama model library for local single-command serving.