WIP-Rec System
An syllabus
Guild to Build a Great Rec System
I have worked at TikTok for two year and half in video recommendation system. TikTok has one of the best recommendation system in the world. I have learned a lot from my colleagues and I want to share my knowledge with you. I want to break down this series in 11 posts:
- Major Component in Rec System
- Modeling Target
- Training Data
- Loss Optimization
- Feature Engineering
- Hyperparameter Tuning
- Model Structure
- Reinforcement Learning in Rec System
- Federated Learning in Rec System
- AutoML and Application in Rec System
- Other
Today we are going to discuss our first topic: Major Component in Rec System.
Major Component in Rec System
In today’s industrial rec system, there are four major components: Recall, Rough Ranking, Fine Ranking, and Mix-Ranker. This design are used by most of the companies in recommendation and ranking system. Of course, some adjustment are made to fit the specific business needs. For example, TikTok user publish millions of videos and our FYP has to select 8 videos out of the pool in real time, thus computation cost and latency is a major concern. However, for some retail E-commerce such as DoorDash or Samsclub.com, the number of viable products are much smaller thus, thus the candidate pool is much smaller, so a rough sort might not be necessary. In this case, a recall + Fine will do the trick just fine.
In this post, I will give introduction to each component, and discuss some design concerns.
Recall
The goal for recall is to select a small subset of candidates from a large pool. This is usually achieved through Multi-Way Recall. Multi-Way Recall builds multiple recall models or rules to select candidates. Some commonly used model targets or rules are:
- Recall recent videos
- Recall popular videos
- Recall Regional Content
- Embedding Based Recall
- I2I: Item 2 Item
- U2I: User 2 Item
- Approximate Nearest Neighbor Recall
Reference
https://zhuanlan.zhihu.com/p/388603950
构建优秀推荐系统指南
我在 TikTok 做了两年半视频推荐系统。TikTok 拥有世界上最好的推荐系统之一。我从同事那里学到了很多,想和你分享这些知识。这个系列拆成 11 篇:
- 推荐系统的主要组件
- 建模目标
- 训练数据
- 损失优化
- 特征工程
- 超参数调优
- 模型结构
- 推荐系统中的强化学习
- 推荐系统中的联邦学习
- AutoML 及其在推荐系统中的应用
- 其他
今天讨论第一个主题:推荐系统的主要组件。
推荐系统的主要组件
当今工业推荐系统有四个主要组件:召回、 粗排、精排 和 混排。大多数做推荐和排序的公司都用这套设计。当然会按具体业务做调整。例如 TikTok 用户每天发数百万视频,FYP 必须实时从池里选出 8 条,计算成本和延迟是主要约束。但对 DoorDash 或 Samsclub.com 这类零售电商,可行商品少得多,候选池也小,粗排往往没必要。这时召回 + 精排就够了。
这篇文章介绍每个组件,并讨论一些设计上的取舍。
召回
召回的目标是从大池里选出一小撮候选。 通常靠 多路召回。 多路召回 用多套召回模型或规则来选候选。 常见目标或规则有:
- 召回最近视频
- 召回热门视频
- 召回地域内容
- 基于 embedding 的召回
- I2I: Item 2 Item
- U2I: User 2 Item
- Approximate Nearest Neighbor 召回
参考
https://zhuanlan.zhihu.com/p/388603950
Comments