Research

Training language models for interviews and learning.

We train our own small models and publish their weights. Two research releases are open on Hugging Face today, both trained on synthetic data with no real child data in them. pinmbo’s safety comes from deterministic guards around whichever model answers.

Where we are

2

open research releases

26M and 4B

parameters in the two models

0

real child records in any training set

Apache-2.0

licence on both

Research directions

Interview-specific training

Genuine fine-tunes for interviewing are rare. We train for follow-up questions, rubric scoring and fair pacing, not general chat.

Beyond Western English

Vernacular and non-Western-English speech is where most interview models fail. It is where we focus.

Integrity signal fusion

Identity, camera, presence, editor and focus signals combined into one picture of a session’s integrity.

Public reliability benchmarks

We believe interview models should be measured in the open, and we are working toward publishing how ours perform.

What the model learns from

Chittu is our model family and the model track behind pinmbo. Two research releases are open on Hugging Face: a 4B guard model trained on 5,250 synthetic examples, and a 26M story model trained from scratch on 2.33 million children’s stories that our own guards swept first. Neither contains real child data and neither serves a child yet.

Our papers and benchmark results will be published on this page.

Our first comparisons are published: How Chittu compares.