Since August 12, 2026, Google has offered a feature on the Pixel 11 capable of translating American Sign Language into English text in real time in Gboard and Live Transcribe. The SL2T model combines computer vision and machine translation, with on-device video processing followed by server-side computation based on body coordinates. This first mainstream integration expands input options, but it does not yet apply to all Android phones or other sign languages, and it is not designed to replace an interpreter in sensitive situations.
What Google Actually Launched
Google DeepMind and Android have integrated Sign-Language-to-Text, or SL2T 1.0, into Gboard and Live Transcribe. The feature converts signs into text in real time. In Gboard, this text can be used in a message, a note, a search, or a query sent to Gemini. In Live Transcribe, the goal is to facilitate informal face-to-face conversations.[1][2]
The initial scope is limited. You need a Pixel 11, the latest version of Gboard, and a keyboard configured for U.S. or Canadian English. The only language pair announced at launch is American Sign Language to English. Google plans to add other devices and languages, but has not provided a specific public timeline.[1][2]
It would therefore be inaccurate to speak of a “generalized” translation of “sign language” on Android. Sign languages are distinct natural languages, each with its own grammar and vocabulary. French Sign Language, for example, is not a visual variant of French and is not supported by this version.
Why This Translation Is More Complex Than a Voice Transcription
Speech recognition generally associates a sound signal with text in a single language. SL2T must perform a translation between two languages. Meaning is conveyed simultaneously through the hands, arms, torso, head, face, space, and rhythm. An approach that converts each sign into an English word would fail to capture part of this structure.
The system therefore directly translates a sequence of movements into text, without resorting to “glosses”—those intermediate annotations often used in research. This approach may better preserve certain spatial and non-manual constructions, but it also makes error analysis more difficult: the model produces a probable sentence without providing a verifiable intermediate linguistic transcription.[1][3]
How SL2T Converts Gestures into Text
The front-facing camera captures the face, upper body, and hands. MediaPipe Holistic, running on the phone, extracts 130 two-dimensional landmarks from each frame. The raw video is then discarded. Only the sequence of coordinates is sent to Google’s servers, where a Transformer-based model generates the English translation.[2][3]
This representation significantly reduces the amount of information conveyed and obscures part of the person’s appearance and surroundings. However, it is not neutral: it omits depth, certain facial details, and articulatory features such as the tongue, as well as objects to which a person might refer. The AISLAC joint report emphasizes that these omissions can affect comprehension.[3]
What the results actually allow us to conclude
Google reports that it trained SL2T on more than 100,000 hours of data covering more than 50 sign languages, about a quarter of which was in ASL. The lab claims that this multilingual training improves transfer between languages and variants. However, the model weights, the full dataset, and the training protocol have not been published, which limits independent reproduction.[1]
On FLEURS-ASL, a studio-recorded ASL-to-English translation dataset, SL2T achieves a BLEURT score of 70 in evaluation without specific adaptation and a BLEU score of 25. The report also mentions a BLEURT score of 74 on a subset signed with a single hand. BLEURT measures the semantic similarity between a translation and a reference, but does not directly represent the reliability of a real-world conversation.[1][3]
The published examples still reveal confusion regarding rare characters, rapid speech, verb tenses, and certain spatial descriptions. As of September 23, 2026, no large-scale independent evaluation is available. The launch therefore demonstrates a useful capability within a defined context, not a universal translation or an equivalent to a professional interpreter.
A step forward in accessibility, provided that users' choices are preserved
The most tangible benefit is a new way to enter text. A signer can produce text without constantly having to use an English keyboard, even if written English is a second language. The tool can reduce daily friction in messaging, searches, and short exchanges.
This utility should not lead to machine translation being imposed as the default interface. Deaf and hard-of-hearing people have diverse language practices. Some prefer sign language, while others prefer writing, speech, captions, or a combination of these. Accessibility therefore requires multiple options, easy correction, and the ability to switch back to the keyboard.
Participatory governance is a key element here. Google established the AI Sign Language Advisory Committee in collaboration with organizations, academic institutions, interpreters, and Deaf experts. The joint report documents the intended uses and limitations. While this consultation does not replace an independent audit, it gives the affected communities a more substantial role than that of a mere testing phase.[3]
Privacy: A Protective Architecture, but Not Entirely Local
The phone does not transmit raw video. This data reduction minimizes risks related to facial images, home locations, and the surrounding environment. Google also states that contact information, videos, and translations are neither stored nor used for training purposes, unless explicitly authorized as part of a study.[2][3]
However, translation remains dependent on the cloud. Physical characteristics are data derived from a person’s behavior and may fall under the GDPR when they relate to an identifiable user. They become sensitive biometric data within the meaning of Article 9 only if they are processed to uniquely identify a person. The legal basis, information, security of the transfer, and the effectiveness of the non-retention policy therefore remain compliance issues that must be documented.[4]
Accessibility and the Law: What the European Framework Entails
The European Accessibility Act has been in effect since June 28, 2025, and applies to various products and services, including smartphones and certain digital services. It imposes functional accessibility requirements without prescribing a specific technical solution such as SL2T. While such a feature may contribute to accessibility, its presence alone is not sufficient to demonstrate a product’s overall compliance.[5]
The European regulation on artificial intelligence does not automatically classify a consumer-grade translation tool as a high-risk system. The context of use remains the determining factor. Using an unevaluated translation in a medical, legal, educational, or professional decision would increase the consequences of an error and could trigger other sector-specific obligations. The explicit limitation to low-stakes situations is therefore essential.[3][6]
What to Watch for Now
- The effective inclusion of other sign languages, with data, assessments, and consultations specific to each community.
- Availability on other Android devices and actual performance may vary depending on the camera, lighting, subject, and framing.
- Independent evaluations of accuracy, latency, serious errors, and differences among groups of signers.
- The possibility of more localized processing, in order to reduce reliance on the network and the exposure of personal data.
- Maintaining access to human interpreters when accuracy, confidentiality, or accountability require it.
SL2T marks an important milestone because a sign language translation model is moving out of the lab and becoming a tool for the general public. Its value, however, will not be measured solely by a benchmark score. It will depend on its reliability in real-world situations, its expansion to multiple languages, data protection, and, above all, its ability to enhance the autonomy of signers without limiting their choices.
Learn more
To learn more about the connections between accessibility, multimodal interfaces, and on-device AI processing, check out these analyses on the aivancity blog.
Sources
[1] Google DeepMind, August 12, 2026, Putting Sign Language AI into Users’ Hands. Read the article
[2] Google Pixel Help, Use Sign-to-Text to Translate American Sign Language into English Text in Real Time, accessed September 23, 2026. View the documentation
[3] Google DeepMind and AISLAC, August 12, 2026, AISLAC Joint Impact Report for SL2T 1.0. View the report
[4] European Union, Regulation (EU) 2016/679 on data protection (GDPR), Articles 4 and 9. View the regulation
[5] European Commission, European Accessibility Act, official fact sheet. View the official fact sheet
[6] European Union, Regulation (EU) 2024/1689 establishing harmonized rules on artificial intelligence (AI Act). View the regulation
