All entriesContact
Model

Soofi

Publisher
Soofi Consortium (Sovereign Open Source Foundation Models) (DE)
Family
Soofi-S
Openness
open_weights
Licence
Permissive open-source committed (OSAID 1.0) but not yet finalised - SPDX pending, LICENSE access-restricted
Context
up to 256k

Soofi-S is an open-science project in the OLMo mould whose pretraining report commits it to the Open Source AI Definition (OSAID 1.0) - a pledge to release weights, intermediate checkpoints, full per-source data accounting, hyperparameters and training+eval code - but as of this grounding that release has not landed: the base weights are in a closed-beta phase ("open model weights coming soon"), the open licence is not yet finalised (the LICENSE file is access-restricted and carries no SPDX), and it remains an early preview/beta with a custom architecture and incomplete safety hardening.

Do you really own it?
Partial
none·limited·partial·substantial·full
Analytical input: AOI B · 73.2/100

Floor-weighted, not averaged. The weakest factor caps the level. Nothing here is weak, but use & modify and data control only reaches moderate - so the substantial bar, strong on both use-and-modify and data-control, is not met, and the level is partial.

The documentation and openness intent here are exceptional - a genuine pretraining report with full per-source data accounting and a stated OSAID 1.0 commitment to release weights, checkpoints, code and data permissively - and the German-sovereign infrastructure is a real data-control draw. But this is preview-stage: the weights are 'coming soon' rather than downloadable, and the licence is unconfirmed (page inconsistent, LICENSE file 401'd, no final SPDX), so use/modify and data-control cannot yet be fully exercised. Treat the openness as promised-and-well-documented rather than delivered, and revisit at the open release once the weights land and the permissive licence is finalised.

1

Use and modify freelyCan you run, modify and adapt it with no gate and no field-of-use trap?

Moderate

The OSAID-compliant permissive licence is a stated future commitment, not yet in force: the licence is unconfirmed/pending (page inconsistent, LICENSE file 401'd, no final SPDX) and the model is a preview whose weights are 'coming soon' rather than generally downloadable - so unconditional commercial use, fine-tuning and continued pre-training cannot yet be fully exercised.

How this scores (AOI sub-dimensions)
Openness3/5how much is released - weights, data, code, licence - and how freelyOpenness here is a well-documented commitment, not yet a delivered release, so it earns the open-weights ceiling (3) rather than the fully-open 5.
Legal4/5how permissive and clean the licence is for real commercial useCommitted permissive OSAID-compliant licensing with unconditional commercial use, an EU/Germany jurisdiction, well under the systemic-risk compute threshold, and a published per-source data inventory that makes the EU AI Act training-content summary directly satisfiable.
2

TransparencyDo you know what it is: weights, training, behaviour, and legible terms?

Strong

Openness rests on a real, detailed pretraining technical report with full per-source data accounting plus a stated OSAID 1.0 commitment - you can already see what it is and how it is built - even though the full weights/checkpoints/data release is a future intention, not yet delivered.

How this scores (AOI sub-dimensions)
Provenance4/5how well we can trace and verify what went into the modelDistributed from the verified Soofi-Project org on Hugging Face in bf16 safetensors (plus first-party GGUF/3-bit quants) with checksums, and because the data accounting, code and intermediate checkpoints are public the pipeline is highly auditable (checklist 6/8).
Governance4/5how accountable and well-documented the publisher isAccountable, well-identified publisher: a named government-funded consortium coordinated by the KI Bundesverband with reputable research members, a detailed pretraining technical report and open GitHub repositories.
3

Doesn't fail youIs it reliable and good enough for the job?

Moderate

Reported as the strongest fully-open model on aggregate English and German benchmarks at only ~3B active params, but an early preview: safety features are documented as incomplete and the numbers come from the report and third-party coverage, not from a generally available release.

How this scores (AOI sub-dimensions)
Performance4/5how capable it is relative to its classIndependent coverage and the pretraining report agree Soofi-S is the strongest fully-open model on aggregate English and German benchmarks - ahead of OLMo 3 32B, Apertus 70B, EuroLLM 22B and Alia 40B - matching dense 14-27B models and winning code aggregates among open base models, at only ~3B active parameters.
Operational4/5how practical it is to run, serve and maintain in productionRuns on mainstream stacks - HF transformers (trust_remote_code), vLLM, llama.cpp/llama-server and Ollama - with first-party GGUF and 3-bit quantizations and a near-constant long-context cache.
Safety3/5whether misuse risks are evaluated and guardrails are providedA post-trained Instruct-Preview variant exists and the L3S Research Centre contributes safety/evaluation frameworks, but the Instruct model card documents several privacy and safety features as still incomplete, this is an early preview, and there is no companion guard model or broad independent red-team battery.
4

Doesn't extract your dataDoes running it keep your knowledge and data yours?

Moderate

EU/Germany-domiciled and trained on sovereign German infrastructure, so once self-hosted your data would stay yours - but with weights not yet downloadable and the licence unconfirmed, that self-hosting cannot yet actually be exercised from a settled grant.

How this scores
Not a scored AOI dimension. For a self-hosted model, data-control is a structural property of running the weights yourself, strong by default unless the model phones home or the licence claws back rights. For a hosted API this factor is the retention + train-on-inputs + residency read, scored from the binding terms.

How the AOI score is computed

The seven dimensions above, each scored 0 to 5, weighted and summed to the 0 to 100 headline. The score is the analytical input behind the ownership verdict, not the verdict itself.

DimensionScoreWeightPoints
Openness3/50.1810.8
Provenance4/50.1612.8
Legal4/50.1612.8
Safety3/50.169.6
Performance4/50.1411.2
Operational4/50.129.6
Governance4/50.086.4
HeadlineB · 73.2/100
Dossier coverageAssess 87%Implement 91%Use 72%Support 50%How complete our four-domain documentation is, a measure of our coverage, not of the model. Each domain links to its page.

Sources

Every rating traces to a primary document. Read means the text was verified; unverified means it is known to exist but has not yet been read.

DocumentWhat it grounds
Technical_reportread2026-07-25
Real arXiv report ("A Sovereign, Open-Source Foundation Model for German and English" / PDF "Soofi S Pretraining Report v1.0"): a ~30B (30B-A3B) hybrid Mamba-2/Transformer MoE trained on ~27T primarily German+English tokens on Deutsche Telekom's Industrial AI Cloud in Munich, with full per-source data accounting (public identifier + raw/effective token counts; the ~1.3% Genios corpus reported in aggregate).
Licenceunverified2026-07-25
Soofi-S-Base exists and is hosted on the verified Soofi-Project org on Hugging Face in bf16 safetensors (custom hybrid Mamba-2/MoE modelling code via trust_remote_code=True), but its licence is unconfirmed/pending: the page characterises it inconsistently ("Other"/TODO vs "closed-beta" and "will be released under a permissive license, without gated access"), terms are described as "not yet finalized", the LICENSE file returned HTTP 401, and no final SPDX has been published.
Model cardunverified2026-07-25
The Soofi-Project Hugging Face org publishes the base plus Instruct/Isar/Rhine preview variants with first-party GGUF and 3-bit quantizations, embedded Jinja chat/tool templates, and model cards noting safety features as incomplete.
Documentationread2026-07-25
The public soofi-project/Soofi-Pretraining GitHub repo (based on the Nvidia Nemotron 3 Nano architecture) hosts the training/evaluation code and per-source data-construction scripts, but its README states "Open model weights coming soon" - the model is not yet generally downloadable (Hugging Face currently hosts a preview/internal checkpoint).
Vendor announcementunverified2026-07-25
A consortium press release presents Soofi as a German public-private consortium project for sovereign industrial AI, trained on Deutsche Telekom's Industrial AI Cloud in Munich, coordinated by the KI Bundesverband and funded by the BMWE.
Third-party analysisunverified2026-07-25
CAIRNE's launch page describes the SOOFI consortium, its members (including CAIRNE Gold member L3S Research Centre) and its sovereign open-source foundation-model mission.
Third-party analysisunverified2026-07-25
Independent technology coverage reports Soofi-S as a leading fully-open model that meets the Open Source AI Definition, publishing weights, intermediate checkpoints, training/eval code and a detailed data inventory.
Third-party analysisunverified2026-07-25
Independent coverage reports Soofi-S is the strongest fully-open model on aggregate English and German benchmarks, ahead of OLMo 3 32B, Apertus 70B, EuroLLM 22B and Alia 40B, and winning code aggregates among open base models.