Contact
Assess · Soofi

Can you own it?

Ownership levelPartialnone·limited·partial·substantial·fullAnalytical input B · 73.2/100

This page is a projection of the one entry record, the Use & modify and Transparency factors that Assess covers. The full verdict is set by all four factors together, floor-weighted so the weakest caps the whole.

Which domain expands which factor
  • AssessUse & modify + Transparency
  • ImplementData control + Doesn't fail you
  • UseDoesn't fail you
  • SupportTransparency

Intended & out-of-scope use

Deploy Soofi-S for industrial and regulated German and English work - technical and regulatory documents, code, and agentic workflows where European data-residency and auditability matter. It is the first model from the German Soofi consortium ("Sovereign Open Source Foundation Models"), coordinated by the KI Bundesverband and funded by the German Federal Ministry for Economic Affairs and Energy (BMWE) under the European IPCEI-CIS programme

  • a ~30B hybrid Mamba-2/Transformer Mixture-of-Experts model (30B total, ~3B active per token) trained on ~27 trillion tokens of primarily German and English.

For adoption, deploy the Instruct-Preview (post-trained) variant; the base checkpoint is built for fine-tuning and is out of scope for customer-facing use. Because this is an early preview/beta, high-stakes, autonomous, or safety-critical decision-making is out of scope without the external control stack described in the Implement domain - treat it as a strong open foundation to pilot, not a hardened production endpoint.

Known limitations, bias & failure modes

  • Early preview/beta. Treat capability, safety and tooling as still stabilising, not final.
  • Incomplete safety hardening. The Instruct-Preview model card documents privacy and safety features as still incomplete, and there is no companion guard/classifier model.
  • Custom architecture. The hybrid Mamba-2/MoE ships custom modelling code and must be loaded with trust_remote_code=True; runtime support is still maturing across serving stacks.
  • Strong-open, not frontier-vs-closed. It is the strongest fully open model on English + German aggregates, but that is a claim relative to open baselines, not the largest closed labs.

The offsetting advantage is transparency: because the full per-source data accounting and code are published, you can inspect what went into the model rather than reasoning about an opaque artifact.

Openness tier & components

Soofi-S is aiming at the OLMo-style fully_open tier, but as of this grounding it has not landed there - so it sits at open_weights with Dimension 1 scored 3, not 5. What is genuinely delivered and read is the documentation: the pretraining report publishes full per-source data accounting (source identifiers, raw and effective token counts, epoch multipliers, and even sources evaluated but excluded) and commits the project to the Open Source AI Definition (OSAID 1.0). But that is a pledge, not a release: the base weights are in a closed-beta phase ("open model weights coming soon" on the soofi-project/Soofi-Pretraining README), the training and evaluation code is not yet released, and the open licence is unconfirmed (no SPDX; the LICENSE file returned HTTP 401 when read). A fully_open/5 rating requires all six components actually Open - downloadable weights under a permissive licence with runnable code - none of which can be verified against a read primary document today. The honest read is promised-and- well-documented, not delivered: revisit and re-score when the OSAID release lands. (A further caveat for later: ~1.3% of Phase 1 tokens, the commercially licensed Genios corpus, are reported in aggregate rather than being redistributable, so even the committed release clears OSAID 1.0 but is only nearly compliant with stricter "every token redistributable" open-data proposals.)

License terms & permitted use

Soofi-S is committed to a permissive, OSAID-compliant open license with unconditional commercial use and no field-of-use restriction - you may use, modify, redistribute and commercialise derivatives. Two honest caveats keep this from being a fully settled Apache-style story: the final license SPDX is not yet published, and the base weights are currently in a closed-beta access phase transitioning to the committed open release. This is why the legal dimension scores 4 rather than 5 - the openness intent and released artifacts are excellent, but the terms are still being finalised.

Supply-chain & provenance

Weights are distributed from the verified Soofi-Project org on Hugging Face in bf16 safetensors (no pickle requirement), accompanied by first-party GGUF and 3-bit quantizations, with per-file checksums. Because the data accounting, training code and intermediate checkpoints are all public, provenance is highly auditable - the checkpoint trust checklist scores 6/8. The two missing controls are cryptographic weight signing (Sigstore/model-signing) and SLSA build attestation, and the base weights sit behind a closed-beta access gate for now, which is why provenance is a 4, not a 5. No incidents or malicious-mirror findings are on record. Pin the exact revision and verify checksums on download.

EU AI Act posture

Soofi-S is a GPAI model but sits well under the 10²⁵-FLOPs systemic-risk threshold (~30B total / ~3B active, ~27T tokens), so no Article 55 regime applies. It is committed to a free/open-source release that would meet OSAID 1.0 and is not monetised, which would qualify it for the Article 53 open-source exemption from the Annex XI/XII technical-documentation duties - but the exemption cannot yet be relied on, because the open licence is not finalised (base weights in closed beta, no SPDX), so there is no confirmed free-and-open licence to point to today. The two obligations that survive - a copyright policy and a public training-content summary - are where Soofi-S is unusually strong: because it publishes full per-source data accounting, the training-content summary can be assembled directly from released artifacts (only the ~1.3% Genios slice is aggregate-only). Combined with an EU/Germany jurisdiction and training on sovereign German infrastructure, this makes Soofi-S one of the cleaner EU-AI-Act postures in the registry. A fine-tuner placing a derivative on the EU market may become a provider for that derivative but inherits an unusually complete upstream package.

Benchmarks & evaluation

OneHill did not run its own benchmarks this session; the picture below is aggregated from the Soofi pretraining technical report and independent coverage (ev-perf, ev-arxiv). The consistent finding: Soofi-S is the strongest fully-open model on aggregate English and German benchmarks, ahead of OLMo 3 32B, Apertus 70B, EuroLLM 22B and Alia 40B, matching dense 14-27B models and winning code aggregates in both languages among open base models - remarkable at only ~3B active parameters. This item is marked partial because exact per-task numbers are still being filled in on the public model cards (some are flagged "to be measured"); treat the leaderboard-style claims as third-party/in-class, not as a frontier-vs-closed claim.

How this scores

The ownership factors this domain covers, drawn from the one entry record.

1

Use and modify freelyCan you run, modify and adapt it with no gate and no field-of-use trap?

Moderate

The OSAID-compliant permissive licence is a stated future commitment, not yet in force: the licence is unconfirmed/pending (page inconsistent, LICENSE file 401'd, no final SPDX) and the model is a preview whose weights are 'coming soon' rather than generally downloadable - so unconditional commercial use, fine-tuning and continued pre-training cannot yet be fully exercised.

How this scores (AOI sub-dimensions)
Openness3/5how much is released - weights, data, code, licence - and how freelyOpenness here is a well-documented commitment, not yet a delivered release, so it earns the open-weights ceiling (3) rather than the fully-open 5.
Legal4/5how permissive and clean the licence is for real commercial useCommitted permissive OSAID-compliant licensing with unconditional commercial use, an EU/Germany jurisdiction, well under the systemic-risk compute threshold, and a published per-source data inventory that makes the EU AI Act training-content summary directly satisfiable.
2

TransparencyDo you know what it is: weights, training, behaviour, and legible terms?

Strong

Openness rests on a real, detailed pretraining technical report with full per-source data accounting plus a stated OSAID 1.0 commitment - you can already see what it is and how it is built - even though the full weights/checkpoints/data release is a future intention, not yet delivered.

How this scores (AOI sub-dimensions)
Provenance4/5how well we can trace and verify what went into the modelDistributed from the verified Soofi-Project org on Hugging Face in bf16 safetensors (plus first-party GGUF/3-bit quants) with checksums, and because the data accounting, code and intermediate checkpoints are public the pipeline is highly auditable (checklist 6/8).
Governance4/5how accountable and well-documented the publisher isAccountable, well-identified publisher: a named government-funded consortium coordinated by the KI Bundesverband with reputable research members, a detailed pretraining technical report and open GitHub repositories.
What this means for adoptionThe documentation and openness intent here are exceptional - a genuine pretraining report with full per-source data accounting and a stated OSAID 1.0 commitment to release weights, checkpoints, code and data permissively - and the German-sovereign infrastructure is a real data-control draw. But this is preview-stage: the weights are 'coming soon' rather than downloadable, and the licence is unconfirmed (page inconsistent, LICENSE file 401'd, no final SPDX), so use/modify and data-control cannot yet be fully exercised. Treat the openness as promised-and-well-documented rather than delivered, and revisit at the open release once the weights land and the permissive licence is finalised.

Sources

The same evidence records as the entry sheet. Read means the text was verified; unverified means it is known to exist but not yet read.

Technical_reportread2026-07-25
Real arXiv report ("A Sovereign, Open-Source Foundation Model for German and English" / PDF "Soofi S Pretraining Report v1.0"): a ~30B (30B-A3B) hybrid Mamba-2/Transformer MoE trained on ~27T primarily German+English tokens on Deutsche Telekom's Industrial AI Cloud in Munich, with full per-source data accounting (public identifier + raw/effective token counts; the ~1.3% Genios corpus reported in aggregate).
Licenceunverified2026-07-25
Soofi-S-Base exists and is hosted on the verified Soofi-Project org on Hugging Face in bf16 safetensors (custom hybrid Mamba-2/MoE modelling code via trust_remote_code=True), but its licence is unconfirmed/pending: the page characterises it inconsistently ("Other"/TODO vs "closed-beta" and "will be released under a permissive license, without gated access"), terms are described as "not yet finalized", the LICENSE file returned HTTP 401, and no final SPDX has been published.
Model cardunverified2026-07-25
The Soofi-Project Hugging Face org publishes the base plus Instruct/Isar/Rhine preview variants with first-party GGUF and 3-bit quantizations, embedded Jinja chat/tool templates, and model cards noting safety features as incomplete.
Documentationread2026-07-25
The public soofi-project/Soofi-Pretraining GitHub repo (based on the Nvidia Nemotron 3 Nano architecture) hosts the training/evaluation code and per-source data-construction scripts, but its README states "Open model weights coming soon" - the model is not yet generally downloadable (Hugging Face currently hosts a preview/internal checkpoint).
Vendor announcementunverified2026-07-25
A consortium press release presents Soofi as a German public-private consortium project for sovereign industrial AI, trained on Deutsche Telekom's Industrial AI Cloud in Munich, coordinated by the KI Bundesverband and funded by the BMWE.
Third-party analysisunverified2026-07-25
CAIRNE's launch page describes the SOOFI consortium, its members (including CAIRNE Gold member L3S Research Centre) and its sovereign open-source foundation-model mission.
Third-party analysisunverified2026-07-25
Independent technology coverage reports Soofi-S as a leading fully-open model that meets the Open Source AI Definition, publishing weights, intermediate checkpoints, training/eval code and a detailed data inventory.
Third-party analysisunverified2026-07-25
Independent coverage reports Soofi-S is the strongest fully-open model on aggregate English and German benchmarks, ahead of OLMo 3 32B, Apertus 70B, EuroLLM 22B and Alia 40B, and winning code aggregates among open base models.