How Chittu compares.
Chittu’s safety does not depend on which model answers. Below is what we measured, against whom, on what date, and what each measurement does not show. Every figure links to the card it came from. Comparisons that have not run yet are listed with their method and no result.
Where Chittu sits
What each product for children in Chittu’s layer says about itself, cited to that vendor’s own page.
Not yet measured
Compared
- pinmbo Chittu
- Six other products for children that occupy the same layer, each named on the card beside its own cited statements
On
- Eight questions asked of every product: where safety lives, open weights, what happens on a crisis message, safety beyond English, consent before use, what a parent sees of the child’s words, engagement mechanics, and stated age range.
How
- Each cell records what the vendor’s own public page says on the date shown, with a link to it.
- Anything a vendor does not say is recorded as “not stated”.
Will report
- A table of public statements, one row per product and one column per question. No score is assigned to anyone.
What this will not show
- This will be a table of what each product says about itself. Nothing in it will be measured by Tulifo.
- “not stated” will mean the vendor’s public materials do not say; it will not be a claim that a feature is absent.
What a child typed and what came back
Whether Chittu’s pipeline contains what the same model lets through without it.
Not yet measured
Compared
- Chittu’s pipeline as deployed Chittu
- The same hosted model without the pipeline, under a kid-safe system prompt published on the card
On
- Our 55-row red-team corpus: 39 attack rows and 16 control rows, 5 passes each.
How
- Same model, same prompts, same day, with and without the pipeline, so the pipeline is the only variable.
- A raw model’s harmful reply is withheld from the public card by rule; the card says what was withheld and why.
Will report
- Attacks contained.
- Over-refusals on the control rows.
- Helplines shown on crisis rows.
- Model calls per turn.
What this will not show
- This will compare a general model under a good prompt with the same model inside the pipeline. It will not be a claim about any vendor.
The story model against TinyStories
Trained only on stories Chittu's guards had already kept, the 26M Chittu story model wrote completions those guards would keep 99.2% of the time on the TinyStories paper's own prompts; the TinyStories models, never filtered this way, came to 89.4–94.7%.
Would pass Chittu's guards (completions) · 131 of 132
Reached an ending within the budget (completions) · 124 of 132
Would pass Chittu's guards (requests) · 60 of 60
Reached an ending within the budget (requests) · 60 of 60
| System | Would pass Chittu's guards (completions) of 132 | Reached an ending within the budget (completions) of 132 | Would pass Chittu's guards (requests) of 60 | Reached an ending within the budget (requests) of 60 |
|---|---|---|---|---|
| Chittu story v0 (26M) Chittu | 99.2% | 93.9% | 100% | 100% |
| TinyStories-1M | 89.4% | 92.4% | — not measured | — not measured |
| TinyStories-3M | 93.9% | 90.9% | — not measured | — not measured |
| TinyStories-8M | 92.4% | 91.7% | — not measured | — not measured |
| TinyStories-33M | 94.7% | 94.7% | — not measured | — not measured |
| TinyStories-Instruct-33M | — not measured | — not measured | 96.7% | 100% |
- Contamination: no evaluation prompt or request appears verbatim in Chittu's training corpus
- expected 0 hits · observed 0 hits over 2,327,256 stories and 64 prompts, whitespace normalised · passed
What this does not show
- A refusal is not a finding of harm. The guards are written for a child's messages and for Chittu's own voice, and they refuse ordinary story lines too. Every refused generation is printed on the full card with the gate that refused it.
- Chittu's model was trained only on stories these same guards had already kept, so this measure favours it by construction. The TinyStories models were never filtered this way.
- The rubric quality of these stories has not been judged by a model or, until the human read is filled, by a person. Hygiene says what the guards would refuse, not whether the story is good.
- No perplexity: no text exists that neither model family saw.
- Different corpora: the TinyStories models trained on the V1 mixed GPT-3.5/GPT-4 data; Chittu on GPT-4 rows swept by its own guards. Sizes differ across the ladder.
- The Instruct model received a translated prompt it was never trained on.
- chittu-story-v0-26m has never served a child and clears none of the gates that would let it.
measured 2026-09-30 · chittu@578729f · full card · Chittu story v0 (26M) on Hugging Face · TinyStories-1M on Hugging Face · TinyStories-3M on Hugging Face · TinyStories-8M on Hugging Face · TinyStories-33M on Hugging Face · TinyStories-Instruct-33M on Hugging Face
The guard model
Whether a small fine-tuned guard flags what general-purpose guard models flag, and what neither has a category for.
Not yet measured
Compared
- Chittu guard v0 (4B) Chittu
- Its base model, Qwen3-4B-Instruct-2507, under the same guard prompt
- Llama Guard 3 (8B)
- ShieldGemma (9B)
- The classifier pinmbo runs in production
On
- The 35-prompt held-out set from the guard model’s card, 3 passes at temperature 0.
How
- Every system sees the same prompts on the same day.
- Scored as flagged or not flagged, because general guards do not emit Chittu’s four tokens.
- A documented mapping from each general guard’s categories to Chittu’s is a secondary table.
Will report
- On guard rows, a flag is correct.
- On safe rows, a flag is an over-refusal.
- Where a general guard has no category for a row, the card reports that absence as the finding.
What this will not show
- The fine-tune track behind this model is paused. It is a research release, not for standalone use.
- A binary flag hides which category a guard chose; the mapping table will show it.