최종 갱신: 2026-08-13
X와 정반대의 자료 상황이다. 코드가 없다. 대신 Transparency Center의 AI system card가 예측 헤드 이름과 개별 피처 문장을 그대로 나열하고, 실적발표(IR) 전사본이 랭킹 변경의 온라인 성과를 분기마다 숫자로 준다. 규제 제출 문서(DSA)와 특허가 블로그보다 구체적인 구간이 있다.
모델·인프라 계보(DLRM → DHEN → Wukong → HSTU → GEM)는 Facebook 전용이 아니므로 topics/로 뺐다. 이 파일은 지면으로서의 Facebook만 다룬다.
system card 인용의 유효기간 — 저장소 공통 주의
이 저장소 전체(services/·topics/)의 Transparency Center system card 인용은 2026-08-09에 직접 읽은 것이다. 다른 파일은 이 절을 링크만 한다. [관찰] 2026-08-09
- 카드 페이지는 client-rendered라
curl로는 본문이 안 나오고, 자동 fetch를 간헐적으로 차단한다 — 2026-08-09 재시도에서curlHTTP 400, WebFetch HTTP 429. 재검증하려면 브라우저 렌더링이 필요하다. - Meta는 카드를 changelog 없이 개정한다. 갱신일만 바뀌고 무엇이 바뀌었는지는 남지 않는다.
- 따라서 카드 기반 주장(예측 헤드 10종, 단계별 후보 수, 신호 50개, 카드 안의 부재 카운트)은 전부 읽은 날짜 기준 스냅샷이다. 인용할 때 카드 갱신일과 이 저장소의 열람일을 같이 적을 것.
한 줄 요약
친구 그래프 조회 + 경량 컷(pass 0) + 멀티태스크 선형 가중합(pass 1) + 컨텍스트 다양성(pass 2)이라는 2021년 구조가 공개된 최신 정본이고, 그 위에서 비연결(추천) 콘텐츠 비중이 2021 11.7% → 2025 Q4 41.0%로 늘어나면서 랭킹 문제가 소셜 그래프 문제에서 순수 추천 문제로 이동했다.
피드 지면
Meta는 지면마다 별도의 system card를 발행한다. 이게 지면 목록의 1차 출처다. [확인] Our approach to explaining ranking, 2023~
- Facebook Feed — 친구/팔로우 페이지/가입 그룹의 연결된 콘텐츠만. 카드 스스로 "Your feed might also include suggested content and advertisements. That content is not powered by the AI we describe in this system card" 라고 범위를 못박는다.
[확인] - Facebook Feed Recommendations — 비연결 콘텐츠. 별도 카드, 별도 예측 헤드 셋.
- Facebook Reels — 세로 단건 소비.
- Facebook Video — 2024년에 Reels/롱폼/Live를 단일 풀스크린 플레이어로 통합.
[확인]Zuckerberg Q2 2024 IR: "This quarter we rolled out our full-screen video player and unified video recommendation service across Facebook -- bringing Reels, longer videos, and Live into a single experience." - Facebook Feed Ranked Comments — 댓글 랭킹이 별도 카드다.
- Friends tab — 2025-03-27 출시, 추천을 빼고 친구 게시물만.
[추정]Zuckerberg 본인 게시물이 1차 출처인데 원문 URL을 확보하지 못했고 보도로만 확인했다.
핵심: 연결/비연결이 두 개의 별도 AI 시스템이고, 최종 피드는 그 둘의 출력에 광고를 끼워 넣은 결과다. [확인] Transparency Center — Our approach to explaining ranking: "a balanced combination of the outputs of both AI systems ... in addition to advertisements and additional product offerings such as groups and reels" (2023-12-31)
⚠ 이 문장은 Nick Clegg 2023 글 본문에 없다 — 2026-08-09에 그 글 전문을 받아 확인. 그 글을 이 인용의 출처로 쓰지 말 것. [관찰] 2026-08-09 (검사한 코퍼스: 해당 URL 본문 1건)
지면 구성 실측 — Widely Viewed Content Report
규제성 공개라 마케팅 문구가 아니다. 미국 오가닉 Feed 조회 비중, 2025 Q4 [확인] WVCR
| 출처 | 비중 |
|---|---|
| Unconnected / Recommended | 41.0% |
| Friends | 20.1% |
| Public Followers | 13.9% |
| Groups | 12.8% |
| Other | 12.2% |
- 링크 포함 조회는 1.5% — "98.5% ... did not include a link to a source outside of Facebook"
[확인] - 추세: 2021 Q3 unconnected 11.7% → 2022 Q3 15.2% → 2025 Q4 41.0%
[확인] - "친구 소식"은 이제 Feed 조회의 5분의 1이다. 이 표 하나가 Meta 랭킹 문헌 전체의 무게중심을 설명한다.
AI 추천 비중 — Meta가 말하다 멈춘 지표
[확인]2023-06-29: "more than 20 percent of content in a person's Facebook and Instagram feeds is now recommended by AI" — ai.meta.com[확인]Q1 2024 Zuckerberg: "about 30% of the posts on Facebook feed are delivered by our AI recommendation system. That's up 2x over the last couple of years. And for the first time ever, more than 50% of the content people see on Instagram is now AI recommended."[관찰]2026-08-09 Q1 2024 이후 이 % 지표는 다시 발표되지 않았다. 검사한 코퍼스는 Q2 2024 ~ Q1 2026 실적발표 전사본 8건 전수이고, 그 안에서 해당 % 표현이 0회다. 이후 지표는 전부 time-spent 증분으로 대체됐다.- → 2026년에 유통되는 "Facebook 피드의 30%가 AI 추천" 류 수치는 2년 묵은 값이다. WVCR의 41.0%(2025 Q4)가 더 최신이고 정의도 다르다(조회 기준 vs 배포 기준). 섞어 쓰지 말 것.
후보 생성 (retrieval)
연결된 콘텐츠에는 ANN이 없다. 코퍼스가 이미 유저의 연결로 제한돼 있어서 임베딩 검색이 필요 없다. 임베딩 retrieval은 전부 비연결 쪽에 몰려 있다. [추정] — 근거: 아래 인벤토리 정의와 pass 0 서술이 그래프 조회 + 경량 모델 컷으로만 구성된다.
인벤토리
[확인] How machine learning powers Facebook's News Feed ranking algorithm, 2021
- 정의: "any non-deleted post shared with Juan by a friend, Group, or Page that he is connected to that was made since his last login" — 마지막 로그인 이후라는 시간 경계가 명시적이다.
- 규모: "more than 2 billion people (more than 1,000 posts per user, per day, on average)"
단계별 후보 수 — system card가 지면마다 다르게 서술한다
[확인] 각 system card
| 지면 | pass 0 (경량 모델) | 최종 스코어링 |
|---|---|---|
| Facebook Feed | 약 500 | 약 500 |
| Facebook Video | 약 1,000 | 상위 200 |
| Facebook Reels | 약 10–100 | (별도 수치 없음) |
| Facebook Feed Recommendations | 수치 없음 | — |
- ⚠ Feed는 500 → 500으로 단조 감소가 없는데 Video는 1,000 → 200으로 2단 축소다. Meta 문서 자체가 지면마다 퍼널 형태를 다르게 서술한다.
- Reels의 10–100은 Feed의 500보다 한 자릿수 작다.
[추정]세로 스와이프 단건 소비라 슬레이트 크기 자체가 작기 때문으로 보이지만 Meta는 이유를 쓰지 않았다. - 2021 블로그의 pass 0 설명: "a lightweight model is run to select approximately 500 of the most relevant posts for Juan that are eligible for ranking" + "This helps us rank fewer stories with high recall in later passes"
[확인]
비연결 콘텐츠 retrieval
[확인]"retrieval systems that take just hundredths of a second to narrow billions of pieces of content down to thousands and then to a few hundred" — ai.meta.com, 2023- 같은 글에 "a novel hierarchical deep neural retrieval architecture" — 이 한 줄이 Andromeda(2024) → HILL/MoNN(2026) 계보의 시작이다. 상세는
topics/embedding-retrieval.md. - ⚠ 충돌: 비연결 system card는 후보 수를 전혀 안 주는데 같은 시스템을 설명하는 블로그는 "billions → thousands → a few hundred"이라고 쓴다. 두 문서가 다른 해상도로 서술한다.
Facebook Feed retrieval의 소스별 후보 수 분해는 no primary source found. [추정] — 검사한 코퍼스는 Transparency Center 시스템 카드 15종, engineering.fb.com·ai.meta.com·about.fb.com의 랭킹 관련 글(위 출처 목록), 실적발표 전사본이고 그 안에 분해가 없다는 뜻이다. 미공개 문서까지 없다는 근거는 아니다. X가 11개 소스를 코드로 열거할 수 있는 것과 대조적이다.
랭킹
공개된 유일한 수식
[확인] engineering.fb.com, 2021 — verbatim 두 개:
Yijt = f(xijt1; xijt2; … xijtC)
Vijt = wijt1·Yijt1 + wijt2·Yijt2 + … + wijtk·Yijtk
i=포스트,j=뷰어(유저),t=시점,k=예측 태스크.[확인]원문이 "for each post i, we estimate Yijt", *"the characteristics of the post X_it toward viewer j at time t"*라고 쓴다. 가중치w가 (i,j,t)로 인덱싱된다 = 포스트별·유저별·시점별로 개인화된다.- 선형 결합을 고른 이유를 본문이 직접 밝힌다: "Any action a person rarely engages in (for instance, a like prediction that's very close to 0) automatically gets a minimal role in ranking, as Yijtk for that event is very low."
- 숫자 weight는 하나도 없다. 심볼릭 수식이다.
가중치 w를 정하는 방법 = 설문
[확인] 같은 글: "the way we take each prediction into account for Juan is based on the actions that people tell us (via surveys) are more meaningful and worth their time" / "People with higher correlation gain more value from that specific event, as long as we make this method incremental and control for potential confounding variables."
→ 설문 신호 자체는 서비스를 가로지르는 기법이라 topics/survey-signals.md로 뺐다.
3-pass 구조
[확인] 같은 글
- pass 0 — 경량 모델이 인벤토리에서 약 500개를 고른다. high recall 목적.
- pass 1 — "the main scoring pass, where each story is scored independently", "we score each post using multitask neural nets". 포인트와이즈다.
- pass 2 — "the contextual pass. Here, contextual features, such as content-type diversity rules, are added to help diversify Juan's News Feed."
- integrity: "certain integrity processes are applied to every post" (2021 블로그, 랭킹 전 위치)
[확인] - ⚠ integrity 적용 지점이 문서마다 다르다.
- 2021 블로그 = 랭킹 전
[확인] - fb-video 카드 = 랭킹 후 — "Certain integrity processes are applied to all results"
[확인] - fb-feed 카드는 단계를 말하지 않는다. 카드의 문장은 "The system also applies certain integrity processes…" 가 전부이고 pass 0·1·2 어디에도 붙이지 않는다.
[확인]fb-feed 카드 [추정]이 저장소는 fb-feed의 그 문장을 pass 0 부근으로 읽어 왔다. 근거는 카드가 이 문장을 후보 축소 서술 옆에 놓는다는 것, 그리고 fb-video만 *"all results"*라고 써서 명시적으로 랭킹 이후를 가리킨다는 대비뿐이다. 카드 자체에 단계 표기는 없으므로 "fb-feed = pass 0"은 우리 해석이다.
- 2021 블로그 = 랭킹 전
- pass 1의 멀티태스크 신경망 구조(MMoE/PLE 계열인지, 타워 수)는 no primary source found.
[추정]— 검사한 코퍼스는 위 출처 목록의 Meta 블로그·시스템 카드 15종이다.
예측 헤드 — Facebook Feed system card
[확인] fb-feed (최종 갱신 2026-06-24) — 10종, verbatim 요지
- scroll-past 여부
- (포맷 선호를 반영한) scroll-past 여부 — 같은 태스크의 별도 모델이 존재한다
- Story In Feed (SIF) skip 여부
- video channel skip 여부
- 'Show more' 클릭
- Story 풀스크린 진입
- like·reaction 목록 클릭
- connection 게시물에 대한 deep/strong intent
- 외부 링크 클릭 + 그 사이트 체류시간 — 외부 체류까지 예측 대상
- connection과 상호작용할 strong intent
부정 예측이 명시적으로 존재한다. Transparency Center 본문이 hide/snooze/unsubscribe·report·angry reaction 예측을 "분포를 낮추는" 방향으로 쓴다고 직접 서술한다. [확인] Our Approach to Facebook Feed Ranking (최종 갱신 2025-06-11)
같은 페이지가 예측을 사용 빈도 3단계로 계층화한다 — Used most frequently / occasionally / less frequently. 이건 X의 ModelWeights 목록에 대응하는 Meta판이고, 가중치 대신 빈도 계층만 공개한 형태다. [확인]
비연결 카드의 예측 헤드는 다르다
체류시간 예측 / share / 'X' 클릭 / 체류시간 증가 / group 가입 / like / 타 플랫폼 비공개 공유 / 스크린샷 / RSVP 클릭 / 'Show more'
- 스크린샷이 명시적 positive다: "List of the owners the user has screenshot their post in the last 180 days"
- 크로스앱 신호가 들어온다: "How many times a user has shared posts on Messenger in the past 1 month", "Owners whose post the user has shared on whatsapp in the last 180 days"
- 위치 신호: "Your approximate distance from the event in a post"
실제 피처 이름 (fb-feed 카드 발췌, 전부 [확인])
- "At which position is the post located on your News Feed?" — 포지션이 피처다
- "Number of times you have seen posts with the same content type (e.g. video, photo, etc) as the current post in the past hour" — 다양성이 재순위 규칙이자 랭킹 피처다
- "Number of times you quickly scroll past the same friend's post in the past 7 days"
- "A weighted count of the number of times a post has been viewed in 2 hours."
- "An embedding to represent the viewer" — 카드에 명시된 유일한 임베딩 피처
- "Quickly predicts how likely you are to interact with a post ... based on how much your friends have engaged with similar content" — 모델 출력을 피처로 재투입(cascade)
- lookback 창이 2시간 / 7일 / 14일 / 28일 / 30일 / 3개월 / 180일로 계층화돼 있다
[추정](근거: 피처 목록에 반복 등장하는 창 길이)
IR이 주는 온라인 성과 (전부 [확인], 규제성 공개)
| 분기 | 발언 |
|---|---|
| Q1 2025 | "a 7% increase in time spent on Facebook, 6% increase on Instagram, and 35% on Threads" |
| Q2 2025 | "a 5% increase in time spent on Facebook and 6% on Instagram just this quarter" |
| Q3 2025 | "5% more time spent on Facebook"; "surfacing twice as many Reels published that day than at the start of the year" |
| Q4 2025 | "a 7% lift in views of organic Feed and video posts on Facebook"; "surfacing over 25% more Reels published that day than the prior quarter" |
| Q1 2026 | "same-day posts now representing more than 30% of recommended Reels on both Instagram and Facebook, more than double the levels one year ago" |
[추정] freshness(same-day 비중)가 3개 분기 연속 별도 지표로 보고된다 → recency가 명시적 랭킹 목표로 승격됐다.
ConnectionMind — LLM이 그래프 경로를 추론하는 랭킹 레이어 (2026-08 논문, 프로덕션 A/B)
⚠ 테제 변경 후보 — 유기 피드 랭킹의 온라인 서빙 경로에 LLM이 들어간다는 Meta 최초의 1차 진술. 상세는 topics/llm-in-recsys.md의 케이스 절에서 다루고, 여기는 지면·파이프라인 위치 관점만 짧게.
[확인] ConnectionMind: Leveraging Social Networks and Large Language Models for Personalized Recommendation at Meta (arXiv 2608.10187), 2026-08-10 — 저자 6명(Haoyu Han은 Michigan State 인턴, 나머지 5명은 Meta Platforms, Menlo Park). Meta 내부 팀명(Core Ranking / GenAI 등)은 소속 각주에 없다.
- 지면 — 논문이 이름을 안 밝힌다. §3은 "large-scale personalized video recommendation in a social platform setting", §6은 *"short-form video platform"*까지만.
[추정]Facebook(Reels 또는 통합 Video 지면)일 가능성이 가장 높다 — 근거 셋: ① 그래프 스키마가 User/Page/Item + User–Group·Group–Item·Page–Item·Item–Item(Co-watch)라 Groups와 Pages가 1급 관계다. Instagram·Threads는 Groups·Pages 스키마가 없다. ② 주 지표가 watch time / video sessions / exposure — 비디오 피드. ③ 2024 Q2 이후 Facebook Reels/롱폼/Live가 통합 video recommendation service로 묶였고, 이 통합 표면이 그래프 시그널을 그대로 받는다. 단 논문이 표면을 확정하지 않았으므로 서비스명 인용 금지 — "Meta short-form video platform"으로 인용할 것. - 파이프라인 위치 — 하이브리드 서빙
[확인]§6.1: 상위 5–10% 헤비 유저는 Llama 3.1-8B-Instruct가 그래프 경로를 온라인에서 직접 추론하고, 나머지 90–95%는 그 LLM이 오프라인에서 발굴한 meta-path로 학습된 Student GNN이 서빙. "For this cohort, requests are routed to the full ConnectionMind LLM" / "we employ an offline distillation pipeline. The LLM processes historical contexts to discover effective meta-paths" — 순수 retrieval도 순수 랭커도 아니고, 후보 스코어링/리랭킹 자리에 붙는 경로 추론기다 - 모델 = LLM이 JSON으로 그래프 경로를 뱉는다
[확인]§4.1: 상태 s_d는 *"current paths, candidate next-hop edges, and compact textual node profiles"*로 직렬화된 프롬프트, LLM은 *"structured action with two components: path expansions and surfaced item–path pairs"*를 낸다. beam search나 constrained decoding은 명시적으로 안 쓰고 format reward로 사후 검증(§4.3). SFT(BFS 최단경로 teacher-forcing) → GRPO(format reward −1/+1, F1 기반 recommendation reward, step-wise shaping, 가중치 α_rec=0.5·α_step=0.3·α_fmt=0.2) - 온라인 A/B — 수천만 유저·다주 지속
[확인]§6.2 Table 2 (프로덕션 대비 상대 리프트): Watch Time +0.43% ± 0.14%, Exposure +0.33% ± 0.08%, Video Sessions +0.22% ± 0.13% - 오프라인 벤치마크는 Delicious·Foursquare (60/20/20) — Meta 데이터가 아니라 소셜 추천 문헌의 공개 벤치. 베이스라인이 GraphRec/DiffNet++/RecDIFF/BIGRec 등 소셜 계열만이고 LightGCN·PinSage·SASRec·HSTU가 없다
[확인]§5.1 - HSTU·GEM·LLaTTE 인용 0건 — 참고문헌에 Meta 내부 시스템(HSTU, GEM, MetaCLIP, LLaTTE, MARM)이 나오지 않는다. Fan et al. 2019(GraphRec) 하나가 Meta 소속 저자의 유일한 인용
[확인]reference list. 같은 회사의 다른 랭킹 라인과 관계를 논문이 설정하지 않는다 — 별개 팀의 병렬 실험선일 수도 있고 내부 정치일 수도 있으나 판정 근거 없음 - 미공개 (열린 질문으로 이동): 그래프 크기(노드·엣지 수)·훈련 데이터 규모·8B LLM의 hot-path latency/QPS/GPU 비용·A/B 가드레일(광고 매출·integrity·negative effects)·per-cohort/country 분해·베이스 GNN 아키텍처 상세
⚠ Meta 자체 문서 간 정합 검사 필요: fb-feed 카드·fb-video 카드·fb-reels 카드 어디에도 "그래프 경로 추론 LLM"이 명시적 예측 헤드나 처리 단계로 등장하지 않는다 [관찰] 2026-08-13 (카드는 2026-08-09에 한 번 렌더링해 읽은 스냅샷 기준). 카드 렌더링을 다시 하지 못한 상태이므로 카드 갱신 여부는 별도 확인 필요.
통합 모델 — 아직 아니다
널리 오해되는 지점이라 못박아 둔다.
[확인]Zuckerberg Q3 2025: "there are three giant transformers that run Facebook, Instagram, and ads recommendations. ... we're also working on combining these three major AI systems into a single unified AI system" → 셋은 별개이고 통합은 진행형 목표다. 타임라인 미제시.[확인]Q2 2024 원형 발언: "I'd like to see us move towards a single, unified recommendation system ... We're not there"[확인]⚠ 가장 헷갈리는 지점: Susan Li Q4 2025 "After seeing strong success with the consolidation of Facebook Feed and video models in the first half of 2025, in Q4 we consolidated models for Facebook Stories and other surfaces into the overall Facebook model. This, along with a series of back-end improvements, drove a 12% increase in ad quality." — 문장에 "Facebook Feed"가 나오지만 성과 지표가 ad quality다. 이건 광고 랭킹 모델의 지면별 통합이지 유기적 Feed 랭킹 통합이 아니다. 그리고 12%는 모델 통합 단독 효과가 아니다 — 원문이 back-end 개선을 함께 원인으로 든다.[확인]유기/광고 플랫폼 공유는 검증 단계: Q4 2025 "We're also going to start validating the use of ads signals and organic content recommendations as we continue to work towards having a more shared platform for organic and ads recommendations over time."[확인]유기 쪽 연구 명칭: "cross-surface foundation recommendation models" (Q2 2025).
재순위·다양성·필터
X와 비교하면 극단적으로 빈약하다. 알고리즘 이름(MMR/DPP), 감쇠 함수, 임계값 N — 전부 없다. 공개 수준의 차이 자체가 기록할 가치가 있다.
- 다양성이 붙는 자리: pass 2 (포인트와이즈 스코어링 뒤 별도 패스)
[확인]2021.[추정]X의 post-selection 배치와 구조적으로 같다. - 유일한 구체 규칙: "tries to ensure that your Feed has a balanced mix of content types. That means, for example, you wouldn't see multiple video posts in a row"
[확인]fb-feed 카드 - 대응 피처가 실재한다 (위 피처 목록의 "same content type ... in the past hour").
- listwise가 존재한다: "pointwise and listwise predictions" + "personalized methodology for delivery frequency control to optimize for long-term user value"
[확인][ai.meta.com, 2023] - 작성자 다양성 규칙, dedup 규칙, 연속 제한 N — no primary source found.
[추정]— 검사한 코퍼스는 시스템 카드 15종 + engineering.fb.com/ai.meta.com 랭킹 글이다. 상세 →topics/feed-diversity.md.
상세 비교는 topics/feed-diversity.md.
데모션 — 2025년에 대폭 삭제됐다
이 저장소 규칙("2019년 논문이 오늘의 프로덕션이라는 보장은 없다")의 교과서적 사례다.
- 2021-09-23 최초 공개, 3대 가치축(직접 피드백 / 고품질·정확성 유인 / 안전한 커뮤니티)으로 그룹핑
[확인]Sharing Our Content Distribution Guidelines. "약 24종"이라는 개수는[추정]— 발표문에는 개수가 없다(3개 그룹명만 실재, 2026-08-09 재검증). 당시 CDG 목록 페이지에서 세어진 것으로 유통되는 수치이며 아카이브 캡처로 재확인 필요. [확인]현재 "Types of content we demote" 페이지(최종 갱신 2025-07-02)에는 4개만 남았다: Clickbait Links / Engagement Bait / Fact-Checked Misinformation / Content Likely Violating Our Community Standards.[관찰]2026-08-09 개별 페이지 404 실측:borderline-content/,low-quality-comments/,links-to-domains-and-pages-with-high-click-gap/, 그리고 인덱스content-distribution-guidelines/자체가 404. 직접 요청해 본 URL 4건에 한정된 관찰이다.[추정]"engagement-bait/,clickbait-links/는 살아있다"는 앞선 기록은 재현되지 않았다. 2026-08-09에/features/approach-to-ranking/types-of-content-we-demote/{engagement-bait,clickbait-links}/와/features/approach-to-ranking/{engagement-bait,clickbait-links}/가 전부 404였다. 두 항목은 현재types-of-content-we-demote/(200) 본문의 섹션으로만 존재하는 것으로 보인다.[확인]원인은 정책 변경이다. Joel Kaplan 2025-01-07: *"we demote too much content that our systems predict might violate our standards. We are in the process of getting rid of most of these demotions."* + "stop demoting fact checked content" — More Speech, Fewer Mistakes[확인]⚠ Meta 자체 문서 간 미해소 불일치: 4개 페이지(2025-07-02 갱신)와 7개 페이지(2025-04-25 갱신)가 동시에 다른 목록을 게시 중이고, 4개 페이지는 여전히 "Fact-Checked Misinformation"을 demote한다고 쓰는데 Kaplan은 이를 중단한다고 했다.[추정]문서에서 사라진 게 시스템에서 사라졌다는 뜻은 아니다. low-quality comment 데모션은 CDG에서 빠졌지만 fb-feed-ranked-comments 카드(2026-04-02)에 예측 태스크로 살아 있다: "How likely a comment appears spammy, abusive, or policy-violating", "How likely the comment is high-quality and contributes positively to the conversation".
borderline content — 곡선과 개입
[추정]Zuckerberg 2018-11-15 "A Blueprint for Content Governance and Enforcement". 원문 URL(facebook.com/notes/…)이 현재 접근 불가, 동시대 인용 재현으로만 확인.- "no matter where we draw the lines for what is allowed, as a piece of content gets close to that line, people will engage with it more on average — even when they tell us afterwards they don't like the content"
- 차트 이름 "Natural Engagement Pattern", 개입은 "penalizing borderline content so it gets less distribution" — 제거가 아니라 억제, 분포 곡선을 line 근처에서 하락하도록 뒤집는다.
[확인]현재 Transparency Center 서술은 훨씬 약화됐고 곡선 언급이 없다: "If content ... doesn't violate the Community Standards, but might still be problematic or otherwise low-quality, Meta may reduce its distribution" (2025-04-25 갱신)
Click-Gap — 메커니즘은 명확, 현 상태 불명
[확인]정의: "Click-Gap looks for domains with a disproportionate number of outbound Facebook clicks compared to their place in the web graph". 웹 그래프에서 in/out 링크가 많은 도메인이 중심, 적으면 주변. 대상은 "succeeding on News Feed in a way that doesn't reflect the authority they've built outside it" — Remove, Reduce, Inform, 2019- 2025년 현재 공개 가이드라인에서 삭제. 운용 중단 여부는 불명 → 열린 질문.
clickbait / engagement bait — 살아있는 두 정의
- Clickbait Links
[확인]: "lure people into clicking ... by creating misleading expectations" — 수법 3종 (a) withholding information (b) sensationalist phrasing ("You won't believe…") (c) punctuation(all caps, 과도한 느낌표) - Engagement Bait
[확인]: "Posts that explicitly request engagement (such as votes, shares, comments, tags, likes, or other reactions) for purposes other than a specific call to action" [확인]2017년 탐지 방식: "Teams ... reviewed and categorized hundreds of thousands of posts" → ML 학습 → 개별 포스트 데모션 + 반복 게시자 페이지 단위 데모션 둘 다 — Fighting Engagement Bait, 2017
정치 콘텐츠 — 억제 레이어를 랭킹 신호로 흡수
[확인] 2025-01-07: "we're going to start treating civic content from people and Pages you follow on Facebook more like any other content in your feed, and we will start ranking ... based on explicit signals (for example, liking a piece of content) and implicit signals (like viewing posts)" + "we will start recommending more political content based on these personalized signals."
→ 2021년의 "civic content 일괄 억제"에서 개인화 랭킹으로 전환. 같은 플립의 Threads 쪽 타임라인은 services/threads.md에 날짜별로 정리했다.
광고 블렌딩
특허가 공식 블로그보다 훨씬 구체적이다. 이 영역에서 유일하게 기계적인 1차 출처다.
[확인] US 10,345,993 B2 "Selecting content items for presentation in a feed based on heights associated with the content items", Meta Platforms, 출원 2015 / 등록 2019 (자매 특허 US 9,729,495)
- 통합 랭킹: "Scores associated with organic news feed stories and scores associated with advertisements are converted into a common unit of measurement, and advertisements and organic news feed stories are together ranked in a single ranking."
- 상단 버퍼: 광고 전에 오가닉 스토리들의 높이 합이 threshold distance를 넘어야 한다.
- 광고 간 간격: "at least a threshold number of organic news feed stories are presented between advertisements" 또는 "a combination of heights of the organic feed stories exceeds a gap distance threshold" — 개수 기준과 픽셀 높이 기준 둘 다로 표현된다.
- position discount: "a predicted decrease in user interaction with a content item based on the position", 그리고 이 할인이 위에 놓인 후보들의 예측 높이 합에 의존한다.
- 높이 예측 자체가 ML 모델이다: 콘텐츠 타입/언어/댓글 수 + 클라이언트 화면 크기/OS/해상도/앱 버전으로 *"the likely height of the content item when presented"*를 예측.
- → Meta의 광고 인터리빙은 슬롯 인덱스가 아니라 픽셀 거리 기반이다. 다른 어느 공개 자료에도 없다.
공식 문서 쪽 유일한 기계적 서술 [확인] Marketing API: Bidding Overview: "Facebook evaluates your bid_strategy, bid_amount, and the probability of acquiring your optimization_goal to calculate an effective bid".
⚠ 널리 도는 Total Value = Bid × Estimated Action Rate × Ad Quality 는 Meta 1차 출처가 없다. [추정] — 검사한 코퍼스는 Marketing API 개발자 문서의 bidding·delivery 페이지와 Transparency Center 광고 관련 페이지다. 2차 출처들이 곱/합을 서로 다르게 쓴다. 수식으로 인용 금지.
⚠ "Ads Allocation in Feed via Constrained Optimization" (KDD 2020)은 LinkedIn 논문이다. Meta 근거로 쓸 수 없다.
ad load는 제품 문서로 공개되지 않는다. [추정] 어닝콜에 "ad-load optimization"이 impression 성장 기여 요인으로 반복 언급되나 수치 분해 없음.
콜드스타트
신규 콘텐츠 — 여기가 훨씬 구체적이다
[확인] Epinet for Content Cold Start (Meta + Stanford, 2024) arXiv:2412.04484
- 대상: Facebook Reels, 일 약 120M 유저
- 콜드스타트 정의가 숫자로 나온다: "videos which have been shown to fewer than 10,000 users"
- 방법: epinet(epistemic NN)으로 불확실성 근사 + Thompson Sampling. user tower / item tower → overarch(base MLP + epinet)
- 위치: "at the proposal stage of recommendation" = 리트리벌 단계. 랭킹·블렌딩 이전.
- 5일 온라인 A/B: impression +17% (전 노출 구간 집계), 저노출 구간에서 like/impression·완주/impression 개선
- 온라인 추천 시스템에 epinet을 적용한 최초 사례라고 주장
부수 장치 [확인]:
- Meta Interest Learner — *"few-shot learning"*으로 신규 콘텐츠를 오디언스에 매칭, "even when there are very few engagements" (2023)
- SilverTorch의 스트리밍 인덱스 갱신이 사실상 신규 콘텐츠 콜드스타트 인프라다 — 모델 publish 주기와 freshness를 분리해 "targeted updates in-place to the specific tensors in the in-memory model". 상세는
topics/model-freshness.md.
신규 유저
Facebook 신규 유저 콜드스타트에 대한 Meta 공개 자료를 찾지 못했다. no primary source found. [추정] — 검사한 코퍼스는 시스템 카드 15종, engineering.fb.com·ai.meta.com·about.fb.com 랭킹 글, Epinet 논문이다. (Instagram 쪽은 one-hop/two-hop 확장이 공개돼 있다 — services/instagram.md)
피드백 신호
dwell time — 정규화 방식까지 공개했다
[확인]2015-06-12 도입. 편향 보정을 명시: "Some people may spend ten seconds on a story because they really enjoy it, while others ... because they have a slow internet connection" → 해법은 개인 내 상대 비교: "if people spend significantly more time on a particular story than the majority of other stories they look at" — Taking Into Account Time Spent on Stories[확인]2016-04-21 외부 아티클 체류시간 예측. 보정 두 가지: "We will not be counting loading time", "looking at the time spent within a threshold so as not to accidentally treat longer articles preferentially" — 길이 편향을 상한 클리핑으로 처리한다. 그리고 그 이전부터 *"clicked on an article and came straight back"*을 clickbait 신호로 쓰고 있었다고 밝힘 — More Articles You Want to Spend Time Viewing
MSI (2018) — 발표문에 숫자가 없다
[확인]Mosseri 2018-01-11: "we will prioritize posts that spark conversations and meaningful interactions", "inspire back-and-forth discussion in the comments". 부작용도 스스로 예고: "Pages may see their reach, video watch time and referral traffic decrease" — Bringing People Closer Together[확인]발표문의 유일한 숫자는 "live videos on average get six times as many interactions as regular videos".[확인]효과는 3년 뒤 공개: "the change led to a decrease of 50 million hours' worth of time spent on Facebook per day" — Clegg, 2021[추정]MSI 포인트 표(like 1 / reaction·reshare 5 / RSVP 15 / significant comment·share 30, 같은 그룹 0.5, 모르는 사람 0.3, angry 1.5→0)는 Meta 1차 출처가 없다. Haugen 유출 문서에 대한 CNN(2021-10-27)·WaPo(2021-10-26)·WSJ(2021-09-15) 보도의 2차 인용이다. 인용할 때 반드시 "유출 문서 2차 보도"라고 붙일 것.[추정]다만 방향성은 Meta 자체 문서와 일치한다 — Transparency Center가 "angry reaction 예측은 분포를 낮춘다"고 직접 쓴다.
유저 컨트롤
| 컨트롤 | Meta의 표현 | 등급 |
|---|---|---|
| Show more / less | "temporarily increase/decrease the ranking score for that post and posts like it" | [확인] 2022-10-05 |
| (이름 변경) | 2025-03-11 업데이트: "At launch, Interested and Not interested were called Show more and Show less" | [확인] |
| Hide | "hide a post so you won't see that post again. This action also helps to minimize similar content" | [확인] fb-feed 카드 |
| Snooze | "temporarily stop seeing posts from a person, Page or group" | [확인] 같은 곳 |
| Favorites | "their posts will be shown higher ... and you'll see their newest posts first" | [확인] 같은 곳 |
- 배수도 지속기간도 공개된 적이 없다. "temporarily"만 있고 시간 단위 없음.
[추정]— 검사한 코퍼스는 fb-feed 카드 본문, 2022-10-05 뉴스룸 글(2025-03-11 갱신본), ai.meta.com Show more/less 글 3건이다(2026-08-09 열람). 미공개 문서까지 없다는 근거는 아니다. - 다만 희소 신호를 일반화하는 방법은 공개됐다
[확인]: 딥러닝으로 user/post 임베딩을 만들고 Like 데이터에서 transfer learning해 Show More/Less를 학습. 의도 모호성("이 사람 덜 보기"인지 "자동차 주제 덜 보기"인지)을 모델이 해석 — The new AI-powered feature designed to improve Feed, 2022
설문 신호
Meta가 가장 깊게 공개한 영역이고 서비스를 가로지르므로 topics/survey-signals.md로 뺐다. Facebook 관련 핵심만: "Is this post worth your time?" (2019 도입), inspirational / topic interest / negative(angry 다수 포스트) 3종 추가(2021-04).
EU DSA — 미국 페이지에 없는 사실 3개
⚠ 아래 1·2는 검증되지 않은 리드다. 삭제하지 않고 등급만 내려 둔다. [추정] — 출처가 Meta 자체 사이트가 아니라 제3자(panoptykon.org)가 호스팅한 PDF이고, Meta 도메인의 원문 텍스트에는 2026-08-09에 도달하지 못했다. 그리고 이 저장소는 PDF에서 뽑은 수치·인용을 신뢰하지 않는다(PDF 요약 도구가 없는 숫자를 만들어낸 전력 — services/threads.md 출처 절의 같은 주의). 재검증 전까지 1·2를 [확인]으로 인용하지 말 것.
[추정] Facebook DSA Systemic Risk Assessment (평가기간 2023-09~2024-08, 2024-11/12 공개), p.27–28
- 주제별로 신호 가중을 차등한다: "We reduce this risk by limiting the role of shares and comments in the distribution of sensitive topics." — 사실이라면 MSI 계열 신호를 주제별로 다르게 쓴다는 유일한 Meta 자체 진술이다.
[추정](같은 문장이 Instagram 쪽 보고서에도 있다고 적어 뒀는데 그쪽도 문서 직링크가 없다 →services/instagram.md) - "we now have 15 recommender 'System Cards'" — 2023년 발표의 22개, 현재 인덱스의 30개(FB 15 + IG 10 + RL 5)와 숫자가 다르다.
[추정]15는 DSA 추천 시스템 한정 카운트로 보인다. - non-profiling 옵션
[확인]2023-08-22: "users will have the option to view Stories and Reels only from people they follow, ranked in chronological order" + 검색은 "based only on the words they enter" — DSA 대응 발표- ⚠ Stories / Reels / Search만 열거하고 Facebook Feed 본체를 명시하지 않는다. DSA Art. 38은 최소 1개의 non-profiling 옵션을 요구한다.
DSA Art. 27이 요구하는 "main parameters" 전용 EU 문서는 별도로 찾지 못했다 — no distinct primary source found. [추정] Meta는 system card 세트 자체를 Art. 27 이행으로 제시하고 있다.
[추정] 이행 실효성 다툼: Amsterdam 지방법원 2025-10-02, 앱 실행마다 알고리즘 피드로 되돌리는 것을 DSA 금지 dark pattern으로 판단, 이행강제금 일 €100,000(상한 €5M). 2025-12 항소심 유지, Meta는 2026-01 대부분 항소 철회. 판결문 원본 미확인, 2차 보도 기반.
이 서비스에서 나온 일반화 노트
topics/multi-task-ranking.md—V = Σ w·Y(덧셈) vs Instagram Explore의 곱셈형. 같은 회사 안 3종 결합 형태topics/candidate-allocation.md— pass 0 후보 수 표, 광고는 별도 슬롯이 아니라 공통 단위 환산 후 단일 랭킹 + 픽셀 간격 제약topics/feed-diversity.md— 다양성을 알고리즘 수준으로 공개하지 않는다는 사실 자체의 기록topics/embedding-retrieval.md— EBR 2020 하이브리드 역인덱스 NN, SilverTorch, RankGraphtopics/survey-signals.md— 설문을 랭킹 신호로 쓰는 방법topics/model-freshness.md— 지면별 재학습 주기topics/sequence-transformer-ranking.md— HSTU / ULTRA-HSTU / LLaTTE
열린 질문
- pass 1의 멀티태스크 신경망 구조 — MMoE/PLE 계열인지, 타워 수는 몇 개인지. 1차 출처 없음.
- pass 2 다양성 규칙의 형태 — 하드 슬롯 제약인지 점수 페널티인지. 없음.
- value model weight
w의 실제 스케일·부호 범위. "hide/report/angry는 분포를 낮춘다"는 방향만 공개. - 연결(500) 결과와 비연결(수백) 결과를 어떤 규칙으로 인터리브하는가. 두 시스템이 별개라고만 밝힘.
- HSTU가 Facebook Feed에 적용됐는가. 논문이 지면을 익명화("a large internet platform")했다. ULTRA-HSTU(2026)가 *"a large-scale production video serving platform that reaches billions of users daily"*로 좁힌 게 현재 최선. (⚠ "video recommendation service"는 Zuckerberg Q2 2024 IR 문구지 ULTRA-HSTU 문구가 아니다.)
- 2025-01 이후 실제로 몇 개 데모션이 제거됐는가. Meta 자체 페이지 두 개가 불일치(4 vs 7).
- Facebook Feed 본체의 DSA non-profiling 옵션 존재 여부와 구현.
- Facebook 신규 유저 콜드스타트.
- Click-Gap이 현재도 운용되는가.
- 2026-03 Rewarding Original Creators on Facebook — 미열람. originality 데모션의 현행 형태.
- ConnectionMind가 붙는 실제 지면 — 논문이 표면을 안 밝힘. 그래프 스키마상 Facebook Reels/Video가 유력하나 확정 못 함.
- ConnectionMind의 서빙 비용·latency — 8B LLM이 헤비 유저 hot path에 들어간다는 진술은 있으나 QPS·p99 지연·GPU 비용·SLO 전무.
- ConnectionMind와 HSTU/GEM/LLaTTE 계보 관계 — 논문에 인용 0건. 같은 팀인지·별개 라인인지 판정 근거 없음.
출처
Meta 공식 — 엔지니어링/뉴스룸
- How machine learning powers Facebook's News Feed ranking algorithm — 2021, 블로그. 수식과 3-pass의 유일한 1차 출처. ⚠ 같은 글이
/core-infra/슬러그로도 잡힌다 - News Feed Ranking in Three Minutes Flat — 2018, 블로그. ⚠
engineering.fb.com/2018/05/22/…경로는 404 - Bringing People Closer Together (MSI) — 2018
- Taking Into Account Time Spent on Stories — 2015
- More Articles You Want to Spend Time Viewing — 2016
- Fighting Engagement Bait on Facebook — 2017
- Remove, Reduce, Inform (Click-Gap) — 2019
- Sharing Our Content Distribution Guidelines — 2021
- Incorporating More Feedback Into News Feed Ranking — 2021
- New Ways to Customize Your Facebook Feed — 2022 (2025-03-11 갱신)
- The new AI-powered feature designed to improve Feed — 2022
- How AI Influences What You See on Facebook and Instagram — 2023
- The AI behind unconnected content recommendations — 2023
- More Speech, Fewer Mistakes — 2025. 데모션 축소의 근거
- You and the Algorithm: It Takes Two to Tango — 2021. ⚠ 기계적 내용 전무. 신호·가중치 근거로 쓸 수 없다
Transparency Center (규제 대응 1차 문서)
- Our Approach to Facebook Feed Ranking — 2025-06-11 갱신
- Our approach to explaining ranking (system card 인덱스)
- Facebook Feed · Feed Recommendations · Reels · Video · Ranked Comments
- Types of content we demote — 2025-07-02 갱신
- Reducing the distribution of problematic content — 2025-04-25 갱신
- Widely Viewed Content Report
- Regulatory and Other Transparency Reports
- ⚠ EU DSA Systemic Risk Assessment — Facebook — 2024, 규제 제출. 제3자(panoptykon.org) 호스팅 PDF이고 Meta 도메인의 원문 텍스트에 도달하지 못했다. 이 저장소는 PDF 추출 수치를 신뢰하지 않으므로 이 문서 기반 주장은 전부
[추정]이다.
IR (규제성 공개 — 랭킹 성과 수치의 1차 출처)
특허·논문
- US 10,345,993 B2 — Selecting content items ... based on heights — 등록 2019, 특허. 광고 인터리빙의 유일한 기계적 출처
- Epinet for Content Cold Start — 2024, 논문 (Meta + Stanford). Facebook Reels 배포
- ConnectionMind: Leveraging Social Networks and Large Language Models for Personalized Recommendation at Meta (arXiv 2608.10187) — 2026-08-10, 논문. Meta 최초의 유기 피드 온라인 서빙 경로 LLM 진술 (Llama 3.1-8B, 헤비 유저 5–10% hot path + Student GNN 90–95%). 지면은 미공개("short-form video platform")
2차 (등급 하향 필수)
- CNN via KVIA — MSI 점수표 — 2021, 유출 문서 보도
- Washington Post — angry emoji 가중 — 2021
- TechCrunch — Zuckerberg 2018 borderline 노트 인용 재현 — 2018