Two weeks of controlled tests showed what recognition demos hide: fluent text can lose the academic detail that changes an answer. Those failures became the architecture.
Student askedDoes a fraction get bigger or smaller when I divide it?
System heard · 93% confidenceDoes a fraction get bigger or smaller when I multiply it?
Sign First decisionRoute withheld · operation needs confirmation
Simulated stress-test row · not an observed student or recognizer measurement
Five findings that changed the product
We did not find a translator. We found the missing layer.
Scroll the field notes. Every number below comes from a locked Sign First artifact; every limitation stays attached.
01 / 05Fluency ≠ fidelity
ASL?AI
0critical tokens recovered
99.5%mean word error
0%abstention rate
ASL-STEM academic evaluation · Vertex video baseline · 2026-08-08
A convincing sentence can still erase the lesson.
On twelve public ASL STEM clips, a general-purpose video model produced fluent fragments without recovering one predeclared critical academic token—and never abstained.
What we observed
The output looked more readable than it was faithful. In education, that is the dangerous failure: the student sees confidence while the operation, quantity, name, or condition disappears.
What it changed
Sign First no longer treats readable English as proof of preserved intent. The handoff stays closed until the meaning-bearing details are visible to the signer.
Evidence boundary
Twelve professional-interpreter research clips are not spontaneous student questions or a population accuracy study.
A celebrated benchmark score can disappear in the classroom.
A public ASL-Citizen checkpoint reproduced 98.44% top-1 accuracy on its official shard, then recovered none of 23 critical tokens on continuous-ASL holdout clips.
What we observed
The model was not fake; the transfer assumption was. Isolated-sign accuracy did not establish safe performance inside continuous academic signing.
What it changed
Recognition is an interchangeable input, not the moat. Every source must earn a bounded role against the academic risks and signer conditions where it will actually operate.
Evidence boundary
This rejects one checkpoint for this workflow. It does not claim that isolated-sign models or ASL-Citizen are broadly ineffective.
The camera called a non-letter pose “F” at 86%.
A real adult-camera test exposed the failure that polished corpus scores missed: unrelated movement and a middle-finger pose produced confident F predictions.
What we observed
The browser model recognized a familiar landmark geometry outside its training distribution. Softmax confidence could not tell us that the pose was not a supported letter.
What it changed
The live path now rejects unsafe topology before classification, requires independent consensus, waits for stable poses, and abstains on unsupported two-hand conditions.
Evidence boundary
The automated safety stack is verified. The fresh seven-probe physical camera certification remains open and the feature is not signer-independent or student-ready.
Review the detail that changes the answer—not every word.
The academic-risk layer recovered every predeclared answer-changing review target in a locked text-pair evaluation while keeping overall signer confirmation mandatory.
What we observed
A full transcript confirmation loop is safe but exhausting. The useful unit is smaller: surface the number, operation, negation, entity, direction, or condition that could change the answer.
What it changed
This became the product contract: detect semantic risk, ask the smallest accessible clarification, preserve the signer’s final authority, then route confirmed intent.
Evidence boundary
This is an internal adversarial text and structural-cue evaluation—not calibrated ASL recognition confidence, child usability, or an educational outcome study.
Corrections can compound without quietly stockpiling children’s video.
A synthetic governed-learning program built a repair suggestion asset, removed revoked lineage on rebuild, and never made the asset eligible to auto-apply or train a visual encoder.
What we observed
The valuable record is not “a child signed something.” It is the governed difference between a candidate and confirmed intent, with purpose, lineage, revocation, and promotion state attached.
What it changed
The future learning moat is a consented repair-evidence system. Ordinary use, personalization, and voluntary contribution remain separate; real-learner visual learning stays disabled.
Evidence boundary
Synthetic technical evidence proves lifecycle controls, not real-student authorization, recognition improvement, or learning effectiveness.
This exact release is the evidence.
Every deployment resets certification on purpose. What you are reading was re-verified for the deployment serving this page—commit ·······—and recorded in an append-only ledger the site cannot edit.
Recertification in progressA new deployment is being re-verified. Completion stays false until every gate passes for this exact commit.
Stakeholder browser audit
awaiting this release
Two ordered production audits
awaiting this release
Operator physical camera report
awaiting the operator
Ledger receipt
issued when all gates pass
Software gates are recorded by automated audits of the live deployment. The physical report is operator-attested evidence with no retained video, landmarks, or candidate text; it is not cryptographic camera attestation, and none of this claims validated student outcomes.
The build ledger
The work is not a model list. It is a sequence of decisions.
Every retained capability below survived a specific gate. Every rejected path remains in the record so the product cannot quietly drift back toward it.
64models cataloged
17datasets cataloged
95structured findings files
0open-domain recognizers approved
01 / 06Rejected
Evidence?Product
0 / 91critical academic terms recovered
Selected decision · use the ledger to inspect what changed
Tested
Twenty locked public academic ASL clips were sent through the direct general-video path.
Learned
The model produced fluent text and never abstained, even when every declared meaning-bearing term was lost.
Changed
Raw recognition can never become an authoritative academic prompt. The signer-confirmed boundary stays closed until meaning is repaired.
Tested
A public isolated-sign checkpoint was reproduced on its own shard, then evaluated on locked continuous-ASL academic evidence.
Learned
The checkpoint was real. The assumption that its benchmark performance transferred to a classroom question was not.
Changed
Every recognizer now has a bounded job, an independent evidence set, and a visible abstention or repair path.
Tested
Sixty-four cataloged models were dispositioned against isolated-sign, continuous-signing, camera, transfer, license, and privacy constraints.
Learned
No accessible model earned open-domain translation. Two narrower components remained useful when their failure surface stayed explicit.
Changed
Static fingerspelling is a bounded working input. Uni-Sign remains a research candidate source that must abstain and hand control to repair.
Tested
The physical laptop camera was challenged with unrelated movement and the unsupported middle-finger pose captured in the observed failure.
Learned
Classifier confidence describes preference among available labels; it does not prove that the input belongs to any supported class.
Changed
Topology rejection, independent consensus, stable dwell, motion rules, and two-hand abstention now run before a letter can be accepted.
Academic harm is concentrated. A smaller decision can protect the task without asking a student to proofread an entire machine transcript forever.
Changed
The product now detects semantic risk and asks the smallest accessible clarification before the confirmed-intent gate can open.
Tested
A purpose-separated synthetic repair asset was built, evaluated, versioned, revoked, and rebuilt without learner video.
Learned
The useful record is the governed difference between candidate and confirmed intent—not an invisible archive of children signing.
Changed
Private use, personalization, and voluntary contribution are separate purposes. Real-learner visual contribution remains disabled.
Evidence boundary: catalog size and test volume show disciplined coverage, not recognition accuracy, Deaf approval, child usability, or learning impact. Those claims remain closed.
Live read-only instrument
Inspect a coherent evidence signature.
The story above explains what the research taught us. This instrument keeps the underlying source class, coverage, thresholds, failures, and privacy invariants inspectable without letting unlike results blend together.
Answer integrity—Answer equivalence within this evidence set
Silent divergence—High-score answer divergence in this set
Mean word error—The flattering field metric
Raw video retained—Invariant · must remain zero
Threshold band · not product approvalNo verdict
Coverage is incomplete. Do not quote or act on these rates.
— / — covered
What the company is building
The durable layer is assurance—not a single recognizer.
Recognition providers will improve and change. Sign First preserves the interaction contract around them: detect meaning-changing risk, return control to the signer, route only confirmed intent, and preserve evidence about where repair was needed.
01
Verified control
Confirmed text gates every AI handoff.
The platform withholds unreviewed recognition from the tutor and connector boundary. The signer-approved meaning is authoritative.
02
Working, bounded input
Static fingerspelling reads 24 letters on-device.
J and Z are excluded because they require motion. The model abstains conservatively, prevents held-pose repeats, and remains unvalidated across independent signers.
03
Research failure retained
Continuous ASL is real—and far from student-ready.
The locked Uni-Sign diagnostic recovered 26.09% of declared critical tokens at 147.71% mean word error. Sign First exposes that weakness and forces repair instead of marketing fluent output.
04
External evidence required
No student outcome claim exists yet.
Age-specific usability, signer-independent recognition, Deaf-led acceptance, and academic outcomes require governed work with people outside this controlled build.
Governed learning concept · technical controls demonstrated with synthetic repair records · no learner-effectiveness claim
Why confirmed repair can compound
Useful learning without turning ordinary use into silent collection.
The near-term asset is not a hidden archive of children signing. It is a governed record of where a candidate differed from confirmed intent—eligible only under a separate purpose and authorization.
01Confirm
Keep the candidate and signer-approved meaning distinct.
02Authorize
Separate private use, personalization, and voluntary contribution.
03Minimize
Build a purpose-specific repair record without raw video.
04Version
Freeze provenance, lineage, eligibility, and revocation state.
05Evaluate
Train offline, test on untouched evidence, and promote or reject.
06Rebuild
Remove revoked lineage and produce a different verifiable asset.
Visual learning remains disabled for real learners.A future visual program needs separate Deaf-led governance, lawful authority, licensed data, retention limits, and revocation that reaches downstream assets.
Read the rows, not only the rate. Each one names what the student asked, what the system heard, and the resulting harm.
No silent divergences in this evidence set.
04
Subject cut
Scope can narrow only where the evidence supports it.
No rows yet for this evidence source.
05
Decision protocol
Thresholds were fixed before the first run. Only complete governed measurement can create a product decision.
Go bandAIR ≥ 90%SDR < 2%
No pilot is authorized by threshold performance alone.
Conditional bandAIR 75–89%
Confirmation remains permanent and scope narrows.
Stop bandAIR < 75% or SDR ≥ 5%
A governed measurement here stops the scoped product path.
06
Confirmed-text handoff funnel
Live serverless sessions recorded in BigQuery. Signer text is encrypted; video and unreviewed recognition never enter the warehouse.
No rows yet for this evidence source.
Public evidence boundary
This register is read-only.
Benchmark writes, model calls, and evidence promotion remain behind the operator boundary. A model, judge, corpus, or seed change requires a fresh signature-bound run.
The public view can refresh its read-only result without creating a run or changing stored evidence.
Warehouse · isolated ASD evidence registerRecognizer · —Tutor · —Judges · — · historical single-judge resultJudge agreement · not available for this signatureCoverage —/—One coherent measurement signature