偏好数据 不靠人工标注也能教会 AI 懂偏好,聊聊 100 万条 UltraFeedback UltraFeedback 是清华团队发布的超大规模 AI … 由Ainexis 2026-07-20 0 阅读更多
评测基准 57 个科目、15908 道题:MMLU 为什么成了大模型的「通用知识体温计」 MMLU(Massive Multitask Languag… 由Ainexis 2026-07-17 0 阅读更多
预训练语料 1.4 万亿 token 里炼出 1560 亿:C4 是怎么把网页「洗」成语料的 C4(Colossal Clean Crawled Corp… 由Ainexis 2026-07-16 0 阅读更多
偏好数据 让 GPT-4 当裁判,18.3 万条偏好对比:Nectar 是怎么给 RLHF 造数据的 Nectar 是 UC Berkeley BAIR 团队(C… 由Ainexis 2026-07-15 0 阅读更多
领域数据 从 2370 亿个网页里捞出 147 亿 token:OpenWebMath 的数学语料炼金术 OpenWebMath 是 2023 年发布的开源数学网页文… 由Ainexis 2026-07-14 0 阅读更多
预训练语料 800GB 开源语料是怎么攒出来的?The Pile 的 22 个子集拆解 The Pile 是 EleutherAI 于 2020 年… 由Ainexis 2026-07-14 0 阅读更多
开源模型 Llama 4 把上下文拉到 1000 万 token,开源模型第一次这么卷 Meta 在 2025 年 4 月发布 Llama 4 系列… 由Ainexis 2026-07-13 0 阅读更多
领域数据 9000 亿 token 的代码语料:The Stack v2 与 Software Heritage 的来龙去脉 The Stack v2 是 BigCode 项目发布的超大… 由Ainexis 2026-07-13 0 阅读更多
指令数据 1.6 万条对话、35 种语言:OpenAssistant 那份众包指令数据的来龙去脉 OASST1(OpenAssistant Conversat… 由Ainexis 2026-07-13 0 阅读更多
偏好数据 「有用」还是「无害」?Anthropic 那套 HH-RLHF 偏好数据怎么教会模型取舍 HH-RLHF 是 Anthropic 在 2022 年论文… 由Ainexis 2026-07-11 0 阅读更多
指令数据 5000 名员工众包 1.5 万条指令:Dolly 那份「能商用」的指令数据怎么来的 当 Alpaca、Vicuna 这些指令微调模型都因为「用了… 由Ainexis 2026-07-09 0 阅读更多
Agent工具 从 LangChain 到 LangGraph,Agent 开发框架走到 1.0 了 LangChain 和 LangGraph 在 2025 年… 由Ainexis 2026-07-08 0 阅读更多