simpo

Reference-free preference alignment, simpler than DPO.

  • Post-Training
  • SimPO
  • Preference Optimization
  • Alignment
  • DPO Alternative
  • Reference-Free
  • LLM Alignment
  • Efficient Training

Declared platforms: linux · macos · windows

Install
npx skills add 'https://github.com/NousResearch/hermes-agent/tree/main/optional-skills/mlops/simpo'
Download bundle ↓
main · 24fd22bScanned 2026-09-15

Contributors

GitHub-linked commit authors for this SKILL.md at the saved revision. Co-authors and history before file renames are not included.

File history ↗
View on GitHub
---name: simpodescription: Reference-free preference alignment, simpler than DPO.version: 1.0.0author: Orchestra Researchlicense: MITdependencies: [torch, transformers, datasets, trl, accelerate]platforms: [linux, macos, windows]metadata:  hermes:    tags: [Post-Training, SimPO, Preference Optimization, Alignment, DPO Alternative, Reference-Free, LLM Alignment, Efficient Training] --- # SimPO - Simple Preference Optimization ## Quick start SimPO is a reference-free preference optimization method that outperforms DPO without needing a reference model. **Installation**:```bash# Create environmentconda create -n simpo python=3.10 && conda activate simpo # Install PyTorch 2.2.2# Visit: https://pytorch.org/get-started/locally/ # Install alignment-handbookgit clone https://github.com/huggingface/alignment-handbook.gitcd alignment-handbookpython -m pip install . # Install Flash Attention 2python -m pip install flash-attn --no-build-isolation``` **Training** (Mistral 7B):```bashACCELERATE_LOG_LEVEL=info accelerate launch \  --config_file accelerate_configs/deepspeed_zero3.yaml \  scripts/run_simpo.py \  training_configs/mistral-7b-base-simpo.yaml``` ## Common workflows ### Workflow 1: Train from base model (Mistral 7B) **Config** (`mistral-7b-base-simpo.yaml`):```yaml# Modelmodel_name_or_path: mistralai/Mistral-7B-v0.1torch_dtype: bfloat16 # Datasetdataset_mixer:  HuggingFaceH4/ultrafeedback_binarized: 1.0dataset_splits:  - train_prefs  - test_prefs # SimPO hyperparametersbeta: 2.0                  # Reward scaling (2.0-10.0)gamma_beta_ratio: 0.5       # Target margin (0-1)loss_type: sigmoid          # sigmoid or hingesft_weight: 0.0             # Optional SFT regularization # Traininglearning_rate: 5e-7         # Critical: 3e-7 to 1e-6num_train_epochs: 1per_device_train_batch_size: 1gradient_accumulation_steps: 8 # Outputoutput_dir: ./outputs/mistral-7b-simpo``` **Launch training**:```bashaccelerate launch --config_file accelerate_configs/deepspeed_zero3.yaml \  scripts/run_simpo.py training_configs/mistral-7b-base-simpo.yaml``` ### Workflow 2: Fine-tune instruct model (Llama 3 8B) **Config** (`llama3-8b-instruct-simpo.yaml`):```yamlmodel_name_or_path: meta-llama/Meta-Llama-3-8B-Instruct dataset_mixer:  argilla/ultrafeedback-binarized-preferences-cleaned: 1.0 beta: 2.5gamma_beta_ratio: 0.5learning_rate: 5e-7sft_weight: 0.1             # Add SFT loss to preserve capabilities num_train_epochs: 1per_device_train_batch_size: 2gradient_accumulation_steps: 4output_dir: ./outputs/llama3-8b-simpo``` **Launch**:```bashaccelerate launch --config_file accelerate_configs/deepspeed_zero3.yaml \  scripts/run_simpo.py training_configs/llama3-8b-instruct-simpo.yaml``` ### Workflow 3: Reasoning-intensive tasks (lower LR) **For math/code tasks**:```yamlmodel_name_or_path: deepseek-ai/deepseek-math-7b-base dataset_mixer:  argilla/distilabel-math-preference-dpo: 1.0 beta: 5.0                   # Higher for stronger signalgamma_beta_ratio: 0.7       # Larger marginlearning_rate: 3e-7         # Lower LR for reasoningsft_weight: 0.0 num_train_epochs: 1per_device_train_batch_size: 1gradient_accumulation_steps: 16``` ## When to use vs alternatives **Use SimPO when**:- Want simpler training than DPO (no reference model)- Have preference data (chosen/rejected pairs)- Need better performance than DPO- Limited compute resources- Single-node training sufficient **Algorithm selection**:- **SimPO**: Simplest, best performance, no reference model- **DPO**: Need reference model baseline, more conservative- **PPO**: Maximum control, need reward model, complex setup- **GRPO**: Memory-efficient RL, no critic **Use alternatives instead**:- **OpenRLHF**: Multi-node distributed training, PPO/GRPO- **TRL**: Need multiple methods in one framework- **DPO**: Established baseline comparison ## Common issues **Issue: Loss divergence** Reduce learning rate:```yamllearning_rate: 3e-7  # Reduce from 5e-7``` Reduce beta:```yamlbeta: 1.0  # Reduce from 2.0``` **Issue: Model forgets capabilities** Add SFT regularization:```yamlsft_weight: 0.1  # Add SFT loss component``` **Issue: Poor preference separation** Increase beta and margin:```yamlbeta: 5.0            # Increase from 2.0gamma_beta_ratio: 0.8  # Increase from 0.5``` **Issue: OOM during training** Reduce batch size:```yamlper_device_train_batch_size: 1gradient_accumulation_steps: 16  # Maintain effective batch``` Enable gradient checkpointing:```yamlgradient_checkpointing: true``` ## Advanced topics **Loss functions**: See [references/loss-functions.md](references/loss-functions.md) for sigmoid vs hinge loss, mathematical formulations, and when to use each. **Hyperparameter tuning**: See [references/hyperparameters.md](references/hyperparameters.md) for beta, gamma, learning rate selection guide, and model-size-specific recommendations. **Dataset preparation**: See [references/datasets.md](references/datasets.md) for preference data formats, quality filtering, and custom dataset creation. ## Hardware requirements - **GPU**: NVIDIA A100/H100 recommended- **VRAM**:  - 7B model: 1× A100 40GB (DeepSpeed ZeRO-3)  - 8B model: 2× A100 40GB  - 70B model: 8× A100 80GB- **Single-node**: DeepSpeed ZeRO-3 sufficient- **Mixed precision**: BF16 recommended **Memory optimization**:- DeepSpeed ZeRO-3 (default config)- Gradient checkpointing- Flash Attention 2 ## Resources - Paper: https://arxiv.org/abs/2405.14734 (NeurIPS 2024)- GitHub: https://github.com/princeton-nlp/SimPO- Models: https://huggingface.co/princeton-nlp- Alignment Handbook: https://github.com/huggingface/alignment-handbook    
Discovery context

Discovered by repository scan. No exact path reference found in the snapshot’s root AGENTS.md.