In brief
Three things to keep straight
- Not every translation difference carries equal academic risk.
- Fluent output can hide a missing or changed task detail.
- The current taxonomy turns a broad safety concern into testable product behavior.
Why average accuracy is the wrong classroom question
A single translation score aggregates many kinds of difference. Education needs a second view: which errors change the requested operation, quantity, condition, entity, direction, or expected answer?12
Google’s own release shows why this distinction matters. Its published examples include “prey” becoming “grey,” dropped classifier information, and tense changes. The company presents these limitations transparently; Sign First asks what each kind of miss would do to an academic task.1
A taxonomy is a test plan, not a universal truth
Sign First currently tracks seven internal risk families: operations, quantities, negation, entities and technical fingerspelling, direction, conditions, and other bounded task constraints. Earlier drafts listed only five and treated the framework as settled. The public evidence register now states the scope and open validation boundary.2
The taxonomy tells the product where to pause and what evidence to collect. It does not establish how frequently those failures occur across ASL users, ages, regions, or subjects.2
The target is the smallest useful repair
Blanket transcript review is a defensible scaffold but a poor future state. The intended experience is to interrupt only when the unresolved detail could change the lesson, show that decision clearly, and let the signer accept, repair, or stop.2
After routing, the answer needs its own boundary. Confirming the question cannot guarantee that an LLM response is relevant, grounded, correct, or appropriate for the learner.2
Questions answered
The short version
Is 90% translation accuracy safe enough for school?
A headline percentage cannot answer that. Safety depends on the evaluation population, task, metric, error distribution, abstention behavior, and consequence of each miss.
Did Google actually confuse 3 + 5 with 3 × 5?
No. That pair is an illustrative Sign First risk scenario, not a reported Google error.
Does the taxonomy prove Sign First works for children?
No. It makes the product claim testable. Deaf-led, age-specific usability and outcome evidence are still required.2
Source register
What this article relies on
- Putting sign language AI into users’ handsGoogle DeepMind ↗
Primary product announcement and stated limitations for SL2T on Pixel 11.
- Evidence registerSign First ↗
Public, bounded record of internal tests, failures, and current product limits.
