LANGUAGE MODEL TRAINING FOR DIRECT PREFERENCE OPTIMIZATION AND IMPROVED LANGUAGE MODEL
Patent №
US 12,694,301
Granted
2026-07-28
Filed 2025
Owner
Intuit Inc.
Lab
—
AI components
0
Assignment
None on record
Dataset
AIPD
Application
19206856
A method for training a language model including executing, a number of times, a training step. The training step includes executing, on a prompt, a preferred output, and a non-preferred output, a reference language model to generate a reference score and a policy score. A loss function includes a combination of the policy score and the reference score, and a hyperparameter that modifies the combination of the policy score and the reference score. The hyperparameter includes a variable term, α, that varies with a number of training steps performed. An updated parameter is generated from the loss. The language model is updated by adjusting an initial parameter of the training language model to the updated parameter. The method also includes returning, after convergence, the updated language model as the improved language model.
Ownership
Intuit Inc.