Federated Learning, Live
Overview
A distributed federated learning system built on Python and Flower that benchmarked FedAvg versus FedProx on non-IID user data from the LEAF suite. A FastAPI backend streams round-by-round training loss and accuracy metrics to a real-time web dashboard over Server-Sent Events (SSE).
Approach
- Engineered the distributed pipeline on authentic non-IID client partitions using real LEAF user data rather than synthetic IID splits.
- Implemented FedProx's proximal penalty term directly into the client execution loop to benchmark mathematical convergence under data heterogeneity.
- Built a resilient Server-Sent Events (SSE) telemetry stream connecting client evaluation nodes directly to the browser UI.


Results
Metrics — sourced from the repo (README / model card / test output)
- FedAvg (Shakespeare, 50 clients, 30 rounds)
- 19.8% → 37.4% accuracy
- FedAvg vs. FedProx (Reddit, 200 clients, 50 rounds)
- 10.96% vs. 8.73% accuracy
- Unit/service tests
- 16 passing
- Telemetry transport
- SSE, not WebSocket
FedProx underperforms here — a real, reported negative result
avoids Hugging Face Spaces' documented proxy WS failures
Demonstrated that FedAvg reached 10.96% next-word accuracy on non-IID text streams, while FedProx variants underperformed due to proximal restriction on initial gradient steps — a negative result on FedProx's own marketed benefit, reported as measured rather than smoothed over. Delivered real-time monitoring infrastructure over SSE, chosen specifically because Hugging Face Spaces' proxy has documented WebSocket failures.