Secure Aggregation for Federated Learning: An ACM CCS 2017 Talk

Added:

Federated Learning
Secure Aggregation
Core Protocol Idea
Handling Dropouts
Privacy Issue
Enhanced Protocol
Security Analysis
Robustness & Efficiency
Performance Results
Future Work

Federated Learning

0:00
Playing Section
  • 1

    Defines federated learning: training a global model on a server using data from distributed devices without uploading raw data.

  • 2

    Explains the core problem: securing the aggregation of individual model updates from clients to prevent the server from seeing them.

Fundamentals of Federated Learning (FL), specifically how decentralized training works and the role of the central coordinating server.
Basic cryptographic concepts including secret sharing schemes (e.g., Shamir's Secret Sharing), symmetric encryption, and Diffie-Hellman key exchange.
An understanding of privacy risks in machine learning, specifically how gradient updates can leak sensitive user data to a curious server.
Basic distributed systems concepts, particularly client-server architecture and the practical challenges of network latency and device dropouts.
Combining Secure Aggregation with Differential Privacy (DP) to achieve mathematically rigorous privacy guarantees against reconstruction attacks.
Advanced and highly scalable Secure Aggregation protocols (such as SecAgg+ or FastSecAgg) that optimize communication overhead and computational complexity.
Robustness against malicious adversaries, focusing on Byzantine fault tolerance and detecting model poisoning attacks within secure communication channels.
Practical deployment of secure federated learning pipelines using production-grade frameworks like TensorFlow Federated (TFF), PySyft, or FATE.
4.5K views74likes22:27@associationofcomputingmach2690Original Release: 2018-02-08

This video introduces a practical secure aggregation protocol for federated learning that enables a server to compute the sum of client model updates without learning individual contributions, addressing privacy concerns in distributed machine learning. The protocol uses Diffie-Hellman key agreement to generate masking vectors between client pairs, combined with threshold secret sharing to handle client dropouts, achieving less than double the bandwidth overhead of naive approaches while maintaining security against honest-but-curious adversaries and tolerating up to one-third of clients dropping out.