这篇博客主要讲解 RLHF 具体训练的框架 (DeepSpeedChat,OpenRLHF,verl) 的具体细节,包括每个框架的整体架构,架构内的各部分细节 (包括逻辑细节和代码细节)。(建议先阅读我之前关于 RLHF 的博客 The Basic Knowledge of RLHF (Reinforce Learning with Human Feedback))
Sampling API Calls: Prompt P(x), sample i from 1 ~ n, filter p_i = p_M(|P(x),x_{1:i-1}) > t_s (up to k) -> I. m API calls, for i in I, using prefix [P(x), x_{1:i-1}, ] until generate -> c_i^1 ~ c_i^m Executing API Calls: for eac...
这篇博客参考了Generative Modeling by Estimating Gradients of the Data Distribution,详细讲述了最近大火的 Diffusion Model 的另一个理解/推理角度: Score-based Generative Model 的数学原理及编程。(ps:建议先看完上述的 Generative Modeling by Estimating Gradients of the Data Distribution 博...
IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing·
This paper proposes a Heterogeneous Network based on Contrastive Learning (HCLNet). HCLNet aims to learn high-level representation from unlabeled PolSAR data for few-shot classification according to multi-features.
It is the first to systematically investigate the effectiveness and underlying mechanisms of activation engineering for mitigating hallucinations in VideoLLMs. And it proposes a temporal-aware activation engineering framework for VideoLL...
The research identifies a critical oversight in existing techniques, which predominantly focus on comparing responses while neglecting valuable latent signals embedded within prompt inputs, and which only focus on preference disparities ...
The rise of reasoning models necessitates large-scale verifiable data, for which programming tasks serve as an ideal source. To address this, we propose a Feedback-Driven Iterative Framework for comprehensive test case construction and r...
International Conference on Learning Representations·
It introduces a Response-conditioned Bradley-Terry (Rc-BT) model that enhances the model's capability in length bias mitigating and length instruction following, through training on the augmented dataset. Furthermore, it proposes the Rc-...
To accurately model the intricate nature of length bias and facilitate more effective bias mitigation, it proposes FiMi-RM (Bias Fitting to Mitigate Length Bias of Reward Model in RLHF), a framework that autonomously learns and corrects ...
This technical report introduces Kimi K3, an open-weight, native multimodal Mixture-of-Experts model with 2.8 trillion total parameters, 104 billion activated parameters, and a one-million-token context window for long-horizon coding, ag...
Master's in Information and Communication Engineering·School of Information Science and Technology, University of Science and Technology of China· – Present
Pursuing graduate study with a research focus on large language models.