Training
Preference tuning (RLHF / DPO)
Edit or adapt a modelFine-tuningReferenceOptimize the model against preference or reward signals (RLHF, DPO, RLAIF) — how agentic and tool-call behaviors like NSC-ACE get shaped.
- Objective
- Edit or adapt a model
- Targets
- Weights
- Format
- Reward / preference optimization
- Granularity
- Full-weight or adapter
- Lifecycle
- QAT / QAD · needs gradients
- Calibration
- Full corpus
- Compression
- None — reshapes behavior, not size
- Quality
- Aligns outputs to preferences; can tax general capability
- Hardware
- Preference dataset + RL/DPO training loop
- Runtimes
A worked Spark recipe for this method hasn't been written yet — it lives here as a reference point in the ontology.