Fitness Tracker Data Platform

Design a data platform for a fitness wearable company (like Fitbit or Garmin) that ingests real time sensor data from 10 million active devices streaming heart rate, steps, sleep stages, GPS coordinates, and SpO2 read...

Design a data platform for a fitness wearable company (like Fitbit or Garmin) that ingests real-time sensor data from 10 million active devices streaming heart rate, steps, sleep stages, GPS coordinates, and SpO2 readings every 5 seconds. The platform must power real-time health alerts (abnormal heart rate, fall detection, dangerous SpO2 drops), generate daily/weekly health reports for users, feed ML models for activity classification, and support longitudinal health studies. The system must handle 2 million events per second at peak, store years of time-series data efficiently, comply with HIPAA regulations, and gracefully handle devices that go offline and sync large backlogs when reconnected. How would you design this end to end?

How to Approach This Problem

How to Approach This Problem Three hard problems make this question unique. Most candidates sketch a Kafka pipeline and call it done. What separates a strong answer is showing you've thought through what actually breaks at scale. 1. Offline sync micro bursts can starve real time health alerts. A wearable offline for 14 days buffers 241,920 readings (~9.7 MB). When it reconnects, it dumps everything in ~24 seconds. If that burst lands on the same Kafka topic as live sensor data, one device's backlog can create consumer lag that delays a cardiac alert for another user. Strong answers isolate offline sync onto a separate topic from the start. 2. One size fits all alert thresholds cause alert fa

Clarifying Questions to Ask the Interviewer

Functional Requirements What sensor types are we ingesting? Heart rate (continuous), step count (aggregated per minute?), sleep stages (classified on device?), GPS (during workouts only?), SpO2 (periodic or continuous)? What is the sampling frequency? Every 5 seconds for HR? Every 1 second during workouts? Every minute for ambient readings? What real time alerts are required? Abnormal heart rate (tachycardia 120 bpm at rest, bradycardia <40 bpm), fall detection (accelerometer spike + no movement), SpO2 drops (<90%), arrhythmia detection (irregular RR intervals)? What latency SLA for health alerts? Sub 10 seconds from sensor reading to push notification? Or best effort? Do devices process any

Envelope Estimation & Capacity Planning

Throughput Math Metric Value Calculation Active wearable devices 10,000,000 Given Sensor reading interval 5 seconds Given (HR, steps, SpO2 per interval) Readings per device per day 17,280 86,400 sec / 5 sec Total events per day 172,800,000,000 10M devices x 17,280 readings Events per second (avg) 2,000,000 172.8B / 86,400 Events per second (peak) ~3,000,000 1.5x peak (morning activity spike, all US time zones awake) Average event payload size ~200 bytes device id, timestamp, HR, steps, SpO2, accel xyz, GPS (nullable) Daily raw ingestion volume ~34.5 TB/day 172.8B x 200 bytes Per second ingestion bandwidth ~400 MB/sec 2M events/sec x 200 bytes Key insight: 2M events/sec at 400 MB/sec is massi

Architecture Walkthrough

Why This Isn't Just a "Store Sensor Data" Problem At first glance, a fitness tracker platform sounds straightforward: collect heart rate, store it, show a graph. But at 10 million devices streaming every 5 seconds, the naive approach (HTTP POST per reading, store in Postgres) collapses instantly. You are dealing with 2 million writes/second of time series data that must simultaneously power sub 10 second health alerts (lives depend on it), daily batch reports, ML model training, and longitudinal health research, all while maintaining HIPAA compliance. The architecture separates four concerns into independent, decoupled stages: 1. Ingestion (IoT protocol): How does data leave the device? 2. S

Component Deep Dive

IoT Gateway / MQTT Broker "The IoT gateway is the front door for 10 million devices maintaining persistent connections. It must handle connection storms (morning wake up, post flight reconnection), validate payloads at line rate, and route to Kafka without becoming a bottleneck." MQTT QoS Levels: Why QoS 1 Is the Sweet Spot: QoS Level Guarantee Round Trips Use Case QoS 0 (at most once) Fire and forget 1 Ambient temperature, non critical sensors QoS 1 (at least once) Guaranteed delivery, possible duplicates 2 (PUBLISH + PUBACK) Heart rate, steps, SpO2, our default QoS 2 (exactly once) Guaranteed delivery, no duplicates 4 (PUBLISH, PUBREC, PUBREL, PUBCOMP) Billing events, medication reminders

Data Modeling

Core Tables Raw Sensor Events (TimescaleDB Hypertable, Hot Storage) Partitioning rationale: 1 hour chunks at 2M events/sec produce ~7.2B rows per hour. TimescaleDB's space partitioning by device id (32 partitions) distributes write load across disk. Compression after 2 hours achieves 10 20x ratio on repetitive sensor data, reducing 15 20 TB to ~1 2 TB. Heart Rate Aggregates (Continuous Aggregate) Why cascading aggregates? Building 1 hour from 5 minute from 1 minute is dramatically more efficient than building each from raw data. The 1 hour rollup scans 12 five minute rows instead of 720 raw rows per device per hour. Activity Sessions (Spark Batch Output) Health Alerts User Devices (Device Re

Loading system design guide...