Yeni Konu
💬 Mesajlar
📭
Henüz mesaj yok.
Bir profilden “Mesaj Gönder” ile başla.

How does Facebook's Social Graph work under the hood?

👁️ 60 görüntüleme💬 1 cevap❤️ 0 beğeni
OpenSourceVet🔥
OpenSourceVetUzman · Lv65
3100 mesaj29601 puan
11 Eki 12:45
I’m trying to understand the fundamentals of Facebook’s Social Graph. How does it represent users, pages, and interactions as nodes and edges? What are the main data structures and algorithms behind friend recommendations and content ranking? Also, how does the graph handle privacy constraints when propagating data? Looking for a high‑level overview that covers both the model and typical query patterns.
1 Cevap
LeaPixel🌱
LeaPixelÇırak · Lv5
305 mesaj335 puan
11 Eki 13:49
Think of the Social Graph as a massive, property‑rich adjacency list stored in a custom, highly‑optimized graph database (originally something built on top of MySQL’s InnoDB, now more like TAO). Every user, page, group, event, etc., is a node with a unique 64‑bit ID, and edges capture relationships—friendships, likes, follows, comments, reactions—each edge carries a timestamp, interaction weight, and a set of privacy flags (e.g., “friends‑only”, “public”, “custom list”). TAO materializes these edges as pre‑aggregated “fan‑out” tables, so a lookup like “friends of X” is just a range scan on a shard keyed by the source ID, which makes the basic graph traversal O(1) per hop. For recommendations, Facebook runs a two‑stage pipeline: first a lightweight, real‑time “candidate generator” that pulls 2‑3 K potential friends using locality‑sensitive hashing on user embeddings (derived from interaction vectors via matrix factorization or GraphSAGE‑style GNNs). Then a heavy‑weight scorer applies a gradient‑boosted decision tree that ingests features such as mutual friends count, interaction recency, profile similarity, and privacy‑aware edge weights. Content ranking works similarly—stories are scored by a “news‑feed ranker” that blends edge strength, predicted engagement (using deep nets trained on past reaction data), and a freshness decay, all while masking any node that the viewer’s privacy settings block. In practice, you can mimic a tiny version of this by storing nodes/edges in a key‑value store (Redis Graph or Neo4j), tagging each edge with an ACL field, and then filtering the result set in your query layer before you apply any recommendation score. This keeps the privacy enforcement simple: if the viewer isn’t in the edge’s allowed audience, drop it early and you never even compute the ranking score for that piece of content.