ROSTER

The 22 arms

12 general decoders, 5 purpose-built safety classifiers, 3 trained encoder classifiers, 2 MLM negative controls. Repository, pinned revision, parameter count, licence, origin and readout family come from the harness's arm registry; on-disk snapshot sizes come from the weight download manifests. 20 of the 22 are candidates and 2 are negative controls. 3 are scored, 11 running and 8 queued, over 85.5 GiB of downloaded weights.

RepositoryPinned revisionClassParametersLicenceOriginGating groupReadoutStatus
answerdotai/ModernBERT-large45bb4654a4d5MLM negative control395,881,664apache-2.0USA/France (Answer.AI/LightOn)ungatedmlm_controlscored
answerdotai/ModernBERT-base8949b909ec90MLM negative control149,655,232apache-2.0USA/France (Answer.AI/LightOn)ungatedmlm_controlscored
meta-llama/Llama-Prompt-Guard-2-86Ma8ded8e697ceTrained encoder classifier278,810,882otherUSA (Meta)promptguard2seqclsqueued
protectai/deberta-v3-base-prompt-injection-v290c9989b1a34Trained encoder classifier184,423,682apache-2.0USA (ProtectAI)ungatedseqclsscored
meta-llama/Llama-Prompt-Guard-2-22M11614a155199Trained encoder classifier70,830,722otherUSA (Meta)promptguard2seqclsqueued
google/gemma-3-4b-it093f9f388b31General decoder4,300,079,472gemmaUSA (Google)gemmaletter3queued
microsoft/Phi-4-mini-instructcfbefacb9925General decoder3,836,021,760mitUSA (Microsoft)ungatedletter3running
ibm-granite/granite-4.0-micro56111ae135dfGeneral decoder3,402,836,480apache-2.0USA (IBM)ungatedletter3running
tiiuae/Falcon3-3B-Instruct411bb94318f9General decoder3,227,655,168other (Falcon LLM licence)UAE (TII)ungatedletter3running
meta-llama/Llama-3.2-3B-Instruct0cb88a4f764bGeneral decoder3,212,749,824llama3.2USA (Meta)llama3.2letter3queued
HuggingFaceTB/SmolLM3-3Ba07cc9a04f16General decoder3,075,098,624apache-2.0France/USA (HuggingFace)ungatedletter3running
HuggingFaceTB/SmolLM2-1.7B-Instruct31b70e2e869aGeneral decoder1,711,376,384apache-2.0France/USA (HuggingFace)ungatedletter3running
tiiuae/Falcon3-1B-Instruct28ba2251970aGeneral decoder1,669,408,768other (Falcon LLM licence)UAE (TII)ungatedletter3running
ibm-granite/granite-4.0-1b6a7381ba1f54General decoder1,631,750,144apache-2.0USA (IBM)ungatedletter3running
allenai/OLMo-2-0425-1B-Instruct48d788eca847General decoder1,484,916,736apache-2.0USA (Ai2)ungatedletter3running
meta-llama/Llama-3.2-1B-Instruct9213176726f5General decoder1,235,814,400llama3.2USA (Meta)llama3.2letter3queued
google/gemma-3-1b-itdcc83ea841abGeneral decoder999,885,952gemmaUSA (Google)gemmaletter3queued
mistralai/Shieldstral-1.0-3B003ec7e2b0baPurpose-built safety classifier3,849,090,048apache-2.0France (Mistral)ungatedshieldstralrunning
ibm-granite/granite-guardian-3.2-3b-a800m3de033d89b49Purpose-built safety classifier3,298,793,472apache-2.0USA (IBM)ungatedletter3running
google/shieldgemma-2bd1dffc9c8c92Purpose-built safety classifier2,614,341,888gemmaUSA (Google)gemmashieldgemmaqueued
ibm-granite/granite-guardian-3.1-2b81145486e85cPurpose-built safety classifier2,533,531,648apache-2.0USA (IBM)ungatedletter3running
meta-llama/Llama-Guard-3-1Bacf7aafa60f0Purpose-built safety classifier1,498,482,688llama3.2USA (Meta)llama3.2llamaguardqueued

Parameter count across the 22 arms

8 of 22 need a licence-accepted token to fetch, across 3 separate acceptance groups. The snapshot column is the size the download wrote to disk.

General decoderPurpose-built safety classifierTrained encoder classifierMLM negative control
0 1 2 3 4 5 gemma-3-4b-it (gated) 4.30 shieldstral-1.0-3b 3.85 phi-4-mini-instruct 3.84 granite-4.0-micro 3.40 granite-guardian-3.2-3b-a800m 3.30 falcon3-3b-instruct 3.23 llama-3.2-3b-instruct (gated) 3.21 smollm3-3b 3.08 shieldgemma-2b (gated) 2.61 granite-guardian-3.1-2b 2.53 smollm2-1.7b-instruct 1.71 falcon3-1b-instruct 1.67 granite-4.0-1b 1.63 llama-guard-3-1b (gated) 1.50 olmo-2-1b-instruct 1.48 llama-3.2-1b-instruct (gated) 1.24 gemma-3-1b-it (gated) 1.00 control-modernbert-large 0.40 prompt-guard-2-86m (gated) 0.28 deberta-v3-prompt-injection-v2 0.18 control-modernbert-base 0.15 prompt-guard-2-22m (gated) 0.07
Table view (every plotted value)
ArmClassParametersSnapshot bytesLicenceOriginGatingReadoutStatus
gemma-3-4b-itGeneral decoder4,300,079,4728,639,634,218gemmaUSA (Google)gemmaletter3queued
shieldstral-1.0-3bPurpose-built safety classifier3,849,090,0487,731,632,731apache-2.0France (Mistral)ungatedshieldstralrunning
phi-4-mini-instructGeneral decoder3,836,021,7607,694,059,344mitUSA (Microsoft)ungatedletter3running
granite-4.0-microGeneral decoder3,402,836,4806,815,498,811apache-2.0USA (IBM)ungatedletter3running
granite-guardian-3.2-3b-a800mPurpose-built safety classifier3,298,793,4726,602,467,360apache-2.0USA (IBM)ungatedletter3running
falcon3-3b-instructGeneral decoder3,227,655,1686,465,509,473otherUAE (TII)ungatedletter3running
llama-3.2-3b-instructGeneral decoder3,212,749,8246,434,752,520llama3.2USA (Meta)llama3.2letter3queued
smollm3-3bGeneral decoder3,075,098,6246,167,868,975apache-2.0France/USA (HuggingFace)ungatedletter3running
shieldgemma-2bPurpose-built safety classifier2,614,341,8885,250,578,613gemmaUSA (Google)gemmashieldgemmaqueued
granite-guardian-3.1-2bPurpose-built safety classifier2,533,531,6485,071,957,207apache-2.0USA (IBM)ungatedletter3running
smollm2-1.7b-instructGeneral decoder1,711,376,3843,426,390,374apache-2.0France/USA (HuggingFace)ungatedletter3running
falcon3-1b-instructGeneral decoder1,669,408,7683,348,994,667otherUAE (TII)ungatedletter3running
granite-4.0-1bGeneral decoder1,631,750,1443,273,295,937apache-2.0USA (IBM)ungatedletter3running
llama-guard-3-1bPurpose-built safety classifier1,498,482,6883,006,170,568llama3.2USA (Meta)llama3.2llamaguardqueued
olmo-2-1b-instructGeneral decoder1,484,916,7362,979,535,483apache-2.0USA (Ai2)ungatedletter3running
llama-3.2-1b-instructGeneral decoder1,235,814,4002,480,847,368llama3.2USA (Meta)llama3.2letter3queued
gemma-3-1b-itGeneral decoder999,885,9522,039,072,537gemmaUSA (Google)gemmaletter3queued
control-modernbert-largeMLM negative control395,881,6641,585,715,090apache-2.0USA/France (Answer.AI/LightOn)ungatedmlm_controlscored
prompt-guard-2-86mTrained encoder classifier278,810,8821,131,677,308otherUSA (Meta)promptguard2seqclsqueued
deberta-v3-prompt-injection-v2Trained encoder classifier184,423,682748,866,413apache-2.0USA (ProtectAI)ungatedseqclsscored
control-modernbert-baseMLM negative control149,655,232600,805,271apache-2.0USA/France (Answer.AI/LightOn)ungatedmlm_controlscored
prompt-guard-2-22mTrained encoder classifier70,830,722292,053,522otherUSA (Meta)promptguard2seqclsqueued
Prompt Guard 2's repo names understate its size: the headline 22M and 86M exclude embeddings. Every value plotted here is read at build time from harness/arms.py and harness/weights_manifest{,2,3,4}.json. The table view lists them all.

Gating

8 of 22 arms need a licence-accepted HuggingFace token, across 3 separate acceptance groups: the Gemma family, Llama 3.2, and a third group covering Llama Prompt Guard 2 at licence other. All three were accepted on 2026-09-23. A 403 is the signal that the token is valid and that repo's group is unaccepted; a 401 means the token itself is not. 14 arms fetch with no acceptance.

Anyone repeating the cohort has to accept three separate licences before 8 of the arms will fetch. Community re-uploads of the gated weights were refused: a mirror launders the provenance the licence column records.

Arms per HuggingFace acceptance group

A 403 rather than a 401 is the signal that the token is valid and that repo's group is unaccepted. Community re-uploads of the gated weights were refused, because a mirror launders the provenance the licence column exists to record.

Fetchable with no acceptanceNeeds a licence-accepted token
0 4 8 12 16 ungated 14 gemma 3 llama3.2 3 prompt guard 2 (licence other) 2
Table view (every plotted value)
GroupArmsAcceptanceMembers
ungated14no acceptance neededgranite-guardian-3.1-2b, granite-guardian-3.2-3b-a800m, granite-4.0-1b, granite-4.0-micro, phi-4-mini-instruct, smollm2-1.7b-instruct, smollm3-3b, olmo-2-1b-instruct, falcon3-1b-instruct, falcon3-3b-instruct, shieldstral-1.0-3b, deberta-v3-prompt-injection-v2, control-modernbert-base, control-modernbert-large
gemma3Google Gemma termsshieldgemma-2b, gemma-3-4b-it, gemma-3-1b-it
llama3.23Meta Llama 3.2 community licencellama-guard-3-1b, llama-3.2-3b-instruct, llama-3.2-1b-instruct
prompt guard 2 (licence other)2a third acceptance group covering Llama Prompt Guard 2 at licence `other`prompt-guard-2-86m, prompt-guard-2-22m
Eight arms behind three groups is a reproducibility cost for anyone repeating this. Every value plotted here is read at build time from pinned/roster.json. The table view lists them all.

Encoder backbones

The encoder arm of the cohort is 3 trained classifiers against 2 untrained MLM controls. Both Prompt Guard 2 sizes are DebertaV2ForSequenceClassification with default LABEL_0 / LABEL_1 heads, so all three trained encoders share the DeBERTa-v2 backbone family. That is three checkpoints inside one family.

An earlier framing held that Prompt Guard 2 supplied an encoder backbone independent of DeBERTa. That was wrong and is withdrawn. The only independent encoder backbone in the cohort is ModernBERT, and both ModernBERT arms are controls. A trained-encoder result from this cohort can be shown to be non-checkpoint-specific inside the DeBERTa-v2 family; it cannot be separated from a DeBERTa-family result.

The evidence is in the download record. The weight manifest for both Prompt Guard 2 sizes records architectures as DebertaV2ForSequenceClassification and max_position_embeddings as 512, which is the same backbone family and the same window as the DeBERTa candidate.

Prompt Guard 2's repo names understate its size. The headline 22M is 70,830,722 parameters and the headline 86M is 278,810,882, because the published figures exclude embeddings.

Backbone family across the five encoder arms

Prompt Guard 2 is DebertaV2ForSequenceClassification at both sizes, so all three trained encoders share the DeBERTa-v2 backbone family. That is three checkpoints inside one family. The only independent encoder backbone in the cohort is ModernBERT, and both ModernBERT arms are controls.

Trained classifierUntrained MLM control
0 1 2 3 4 DebertaV2, trained classifiers 3 ModernBERT, untrained controls 2
Table view (every plotted value)
ArmBackboneClassParametersHow the backbone is knownStatus
deberta-v3-prompt-injection-v2DebertaV2Trained encoder classifier184,423,682the repo is a DeBERTa-v3 base checkpoint with a trained 2-class head; the download that recorded it predates the architecture field in the manifestscored
control-modernbert-baseModernBERTMLM negative control149,655,232ModernBertForMaskedLM, recorded in the scorecard's readout as mlm_controlscored
control-modernbert-largeModernBERTMLM negative control395,881,664ModernBertForMaskedLM, recorded in the scorecard's readout as mlm_controlscored
prompt-guard-2-86mDebertaV2Trained encoder classifier278,810,882DebertaV2ForSequenceClassificationqueued
prompt-guard-2-22mDebertaV2Trained encoder classifier70,830,722DebertaV2ForSequenceClassificationqueued
A trained-encoder result here can be shown to be non-checkpoint-specific within the DeBERTa-v2 family. It cannot be separated from a DeBERTa-family result. Every value plotted here is read at build time from pinned/roster.json. The table view lists them all.

Taxonomy coverage in the safety classifiers

Three of the 5 purpose-built safety classifiers carry a policy surface. One has a mismatch with this task.

ArmPolicy mechanismCovers destructive tool calls
meta-llama/Llama-Guard-3-1Bthe shipped template hardcodes 13 categories, S1 to S13no; llamaguard_default_taxonomy_covers_task: false is recorded in the run metadata
google/shieldgemma-2bthe chat template takes a guideline argumentyes, once the I3 policy is passed as the guideline
mistralai/Shieldstral-1.0-3Bpolicy argumentyes, once the I3 policy is passed as the policy

Llama Guard 3 1B's default categories

The 8B model in the same family carries an S14 Code Interpreter Abuse category. The 1B's default list does not. S2 Non-Violent Crimes is the nearest fit and mapping onto it would have been a manufactured mapping, so a single custom category built from the I3 policy plus the Q2 block criterion, carried through the shipped template's documented `categories` hook and recorded verbatim in the run metadata. Readout is P(unsafe) against P(safe) over first-token ids [20451, 39257] against [6220, 19193].

ShieldGemma's chat template takes a guideline argument, so it is policy-adaptable and closer to Shieldstral on this axis than expected. The I3 policy and the Q2 block criterion were passed as the guideline and recorded verbatim, reading P(Yes) against P(No). Shieldstral takes a policy argument and was treated the same way. Llama Guard 3 1B is the only arm in the cohort with the taxonomy mismatch.

Qwen3

Qwen3 at 0.6B, 1.7B and 4B was dropped on a non-China provenance constraint. The download record for those 3 repositories totals 13,659,595,858 bytes, 12.72 GiB, and no GPU time was spent. They were the strongest ungated general models at their sizes.

With Gemma 3 and Llama 3.2 both gated, the ungated non-China general field in this cohort contains nothing that is simultaneously best-in-class for its size and permissively licensed.

Licence column

Both Falcon3 arms are licence other, the Falcon LLM licence, and neither is Apache-2.0. Prompt Guard 2 is licence other at both sizes. The table records each arm's licence as its publisher declares it.

Source: benchmarks/EXPERIMENTS.md on branch feat/system-one-benchmarks, staged into pinned/roster.json and cross-checked by conform_to_space_contract.py. See Reproduce.