Google Research unveils TEE-based federated learning with verifiable privacy

On October 2, 2026, Katharine Daly (Software Engineer) and Daniel Ramage (Research Director) of Google Research announced the next generation of Google's Federated Learning (FL) system. It uses Trusted Execution Environments (TEEs) on the server to give externally verifiable and auditable data anonymization guarantees, while shifting computation to the server to improve training speed, accuracy and device coverage. The accompanying whitepaper is titled "Toward provably private learning from federated data".
Background: Google introduced FL in 2017 as a technique that trains models across decentralized, private data. It powers next-word prediction and Smart Compose on Gboard, reply suggestions in Google Messages and Smart Text Selection in Android. Google says its FL development follows four privacy principles: data minimization, data anonymization, transparency and control, and verifiability and auditability. Years of anonymization research produced strong differential privacy (DP) guarantees for production models, through algorithms such as matrix factorization DP-FTRL (MF-DP-FTRL) and distributed DP coupled with Secure Aggregation. In 2025 Google introduced an evolved definition of FL centered on those four principles.
The problem the new system targets: in earlier FL systems, device data was uploaded for immediate aggregation, but external observers could not verify that it was never logged or inspected. Secure Aggregation later protected uploads cryptographically but was not compatible with state-of-the-art central DP guarantees. The logic on the server could be verified neither by devices nor by auditors, so Google had to be trusted to add random noise to gradient sums correctly. The new TEE-based system is described as the next milestone in an effort to remove the need to trust the server operator. It builds on Google's earlier work on confidential federated analytics and provably private insights.
How it works: logic running in a TEE is remotely attestable (third parties can verify what is being executed), and it gains confidentiality and integrity, subject to current-generation TEE limitations. In the new system, only metrics and differentially private model weights are visible to workload operators. Encrypted training data from devices can be decrypted and processed only inside TEEs running Python training programs listed in access policies, and only for a limited time after upload. Those access policies, which directly describe the Python program expressing the FL training logic, are published to Rekor, a public transparency log. Devices therefore know the full set of server workloads that may access their uploads, and external auditors can watch the log. The key management service (KMS) and data processing binaries can be reproducibly built from open source code in the Confidential Federated Compute Github repository. To protect proprietary model architectures and data preprocessing logic, the data processing TEEs support sideloading serialized information into the Python program at runtime. This works as long as all privacy-relevant logic stays hardcoded in the Python program.
Results so far: Gboard has deployed the system to launch English and Japanese next word prediction models with stronger privacy guarantees and improved accuracy, and reports substantially faster compute times than the previous FL system. Because all device uploads are collected before the server-side training workload runs, diurnal variations in device availability no longer affect training progress, and the program can calculate the optimal device participation schedule dynamically and use it to tune other DP parameters. In the past, training such FL models could take 1-2 months each, limited by device availability, on-device compute and competition among training workloads for the same device resources. Now bottlenecks have moved to the server, and parallelization across many machines gives significant speedups, currently limited only by TEE resource availability. Client gradient computation also moves to the server, lifting on-device compute limits and, the authors say, paving the way for training increasingly larger models with FL. Integrating TEEs with accelerators will matter for that.
What comes next: the system can run verifiably not only FL training but any workload expressible in Python. Google is experimenting with other workloads such as synthetic data generation, and exploring combining these TEEs with other data processing TEEs specialized in functions like LLM inference. The authors call the work a step toward rigorous proof that server-side processing preserves individual privacy. They expect future TEE hardware and research on mitigating side-channel observations to give deeper protection for dynamically loaded workloads against malicious server-side attacks, and anticipate that systems like theirs may one day come with full proofs of correctness for the software implementations of the DP algorithms and system components.
Key facts
- Google Research's new Federated Learning system runs server-side training in Trusted Execution Environments, so third parties can verify the code that processes device data.
- Only metrics and differentially private model weights are visible to workload operators; access policies are published to the Rekor public transparency log.
- Gboard has deployed it for English and Japanese next word prediction, with improved accuracy and substantially faster compute than the previous FL system.
- Earlier FL models could take 1-2 months each to train; the new system moves bottlenecks to the server and parallelizes across machines, limited only by TEE resource availability.
- The authors call it a step toward, not yet, rigorous proof of server-side privacy, and acknowledge current-generation TEE limitations and open side-channel questions.
Why it matters
Federated learning has always promised that raw data stays private, but the server side was a matter of trust: neither devices nor auditors could check that Google added DP noise correctly or that uploads were never logged. This system aims to replace that trust with verification, using remote attestation, published access policies and open source binaries. It also removes a practical cost. Training that once took 1-2 months per model, limited by device availability and on-device compute, now runs on parallelized server machines.
Who it affects
Users of Gboard are the first in line, since English and Japanese next word prediction models already run on the new system. The announcement also names other FL-powered features, including Smart Compose, reply suggestions in Google Messages and Smart Text Selection in Android, as products of the FL approach generally. External auditors gain a public log (Rekor) to inspect. Privacy and ML researchers get an open source codebase, the Confidential Federated Compute Github repository, and a whitepaper.
How to use it
This is Google's internal production infrastructure, not a product for outside use, and the post gives no price or availability for anyone else. What outsiders can do is inspect it: read the whitepaper "Toward provably private learning from federated data", build the KMS and data processing binaries from the Confidential Federated Compute Github repository, and watch the access policies published to the Rekor transparency log.
How solid is it
This is Google Research's own announcement, written by two Google Research authors, and the claims are the authors' own. The post reports Gboard's deployment but no numeric accuracy gain for the English and Japanese models, no exact training time for the new system, and no differential privacy epsilon or other numeric privacy parameters. No independent audit or third-party evaluation results are reported. The authors themselves frame it as a step toward rigorous proof, not the proof.
Risks and caveats
The confidentiality and integrity guarantees hold subject to current-generation TEE limitations, and the authors say they expect future TEE hardware and ongoing research on side-channel observations to give deeper protection for dynamically loaded workloads against malicious server-side attacks. The sideloading feature keeps proprietary logic hidden, which works only as long as all privacy-relevant logic is hardcoded in the Python program. The post does not say the system eliminates all risk. No hardware vendor or TEE technology is named, and there is no timeline for rollout beyond Gboard or for the experimental workloads such as synthetic data generation and LLM inference.
“This work is a step toward rigorous proof that server side processing preserves individual privacy.”
— Katharine Daly and Daniel Ramage, Google Research