guidance

Constrain LLM output with grammars; guarantee valid JSON.

  • Prompt Engineering
  • Guidance
  • Constrained Generation
  • Structured Output
  • JSON Validation
  • Grammar
  • Microsoft Research
  • Format Enforcement
  • Multi-Step Workflows

Declared platforms: linux · macos · windows

Install
npx skills add 'https://github.com/NousResearch/hermes-agent/tree/main/optional-skills/mlops/guidance'
Download bundle ↓
main · 24fd22bScanned 2026-09-15

Contributors

GitHub-linked commit authors for this SKILL.md at the saved revision. Co-authors and history before file renames are not included.

File history ↗
View on GitHub
---name: guidancedescription: Constrain LLM output with grammars; guarantee valid JSON.version: 1.0.1author: Orchestra Researchlicense: MITdependencies: [guidance, transformers]platforms: [linux, macos, windows]metadata:  hermes:    tags: [Prompt Engineering, Guidance, Constrained Generation, Structured Output, JSON Validation, Grammar, Microsoft Research, Format Enforcement, Multi-Step Workflows] --- # Guidance: Constrained LLM Generation ## When to Use This Skill Use Guidance when you need to:- **Control LLM output syntax** with regex or grammars- **Guarantee valid JSON/XML/code** generation- **Reduce latency** vs traditional prompting approaches- **Enforce structured formats** (dates, emails, IDs, etc.)- **Build multi-step workflows** with Pythonic control flow- **Prevent invalid outputs** through grammatical constraints **GitHub Stars**: 18,000+ | **From**: Microsoft Research ## Installation ```bash# Base installationpip install guidance # With specific backendspip install guidance[transformers]  # Hugging Face modelspip install guidance[llama_cpp]     # llama.cpp models``` ## Quick Start ### Basic Example: Structured Generation ```pythonfrom guidance import models, gen # Load model (supports OpenAI, Transformers, llama.cpp)lm = models.OpenAI("gpt-4") # Generate with constraintsresult = lm + "The capital of France is " + gen("capital", max_tokens=5) print(result["capital"])  # "Paris"``` ### Chat format with a local model > **Constraint support requires local logit access.** Regex, `select()`, and> grammar-based constrained generation only work with local backends> (`Transformers`, `LlamaCpp`). Remote API backends (`OpenAI`, and Azure> variants) support unconstrained `gen()` / chat only — they cannot enforce> token-level constraints. guidance 0.3.x has no `models.Anthropic` class. ```pythonfrom guidance import models, gen, system, user, assistant # Local model (supports constrained generation)lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Use context managers for chat formatwith system():    lm += "You are a helpful assistant." with user():    lm += "What is the capital of France?" with assistant():    lm += gen(max_tokens=20)``` ## Core Concepts ### 1. Context Managers Guidance uses Pythonic context managers for chat-style interactions. ```pythonfrom guidance import system, user, assistant, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # System messagewith system():    lm += "You are a JSON generation expert." # User messagewith user():    lm += "Generate a person object with name and age." # Assistant responsewith assistant():    lm += gen("response", max_tokens=100) print(lm["response"])``` **Benefits:**- Natural chat flow- Clear role separation- Easy to read and maintain ### 2. Constrained Generation Guidance ensures outputs match specified patterns using regex or grammars. #### Regex Constraints ```pythonfrom guidance import models, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Constrain to valid email formatlm += "Email: " + gen("email", regex=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") # Constrain to date format (YYYY-MM-DD)lm += "Date: " + gen("date", regex=r"\d{4}-\d{2}-\d{2}") # Constrain to phone numberlm += "Phone: " + gen("phone", regex=r"\d{3}-\d{3}-\d{4}") print(lm["email"])  # Guaranteed valid emailprint(lm["date"])   # Guaranteed YYYY-MM-DD format``` **How it works:**- Regex converted to grammar at token level- Invalid tokens filtered during generation- Model can only produce matching outputs #### Selection Constraints ```pythonfrom guidance import models, gen, select lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Constrain to specific choiceslm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment") # Multiple-choice selectionlm += "Best answer: " + select(    ["A) Paris", "B) London", "C) Berlin", "D) Madrid"],    name="answer") print(lm["sentiment"])  # One of: positive, negative, neutralprint(lm["answer"])     # One of: A, B, C, or D``` ### 3. Token Healing Guidance automatically "heals" token boundaries between prompt and generation. **Problem:** Tokenization creates unnatural boundaries. ```python# Without token healingprompt = "The capital of France is "# Last token: " is "# First generated token might be " Par" (with leading space)# Result: "The capital of France is  Paris" (double space!)``` **Solution:** Guidance backs up one token and regenerates. ```pythonfrom guidance import models, gen lm = models.Transformers("microsoft/Phi-4-mini-instruct") # Token healing enabled by defaultlm += "The capital of France is " + gen("capital", max_tokens=5)# Result: "The capital of France is Paris" (correct spacing)``` **Benefits:**- Natural text boundaries- No awkward spacing issues- Better model performance (sees natural token sequences) ### 4. Grammar-Based Generation Define complex structures by composing grammar functions. The template-string`grammar=` form is not part of current guidance — build grammars fromcomposable functions, or use `guidance.json()` for JSON. ```pythonfrom guidance import models, genfrom guidance import json as gen_jsonfrom pydantic import BaseModel, Field lm = models.Transformers("microsoft/Phi-4-mini-instruct") # JSON via a Pydantic schema (guidance.json compiles the schema to a grammar)class Person(BaseModel):    name: str = Field(pattern=r"[A-Za-z ]+")    age: int    email: str = Field(pattern=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") lm += gen_json(name="person", schema=Person) print(lm["person"])  # Guaranteed valid JSON matching the schema # Or compose grammar functions directly:grammar = "name=" + gen("name", regex=r"[A-Za-z ]+") + " age=" + gen("age", regex=r"[0-9]+")lm += grammar``` **Use cases:**- Complex structured outputs- Nested data structures- Programming language syntax- Domain-specific languages ### 5. Guidance Functions Create reusable generation patterns with the `@guidance` decorator. ```pythonfrom guidance import guidance, gen, models @guidancedef generate_person(lm):    """Generate a person with name and age."""    lm += "Name: " + gen("name", max_tokens=20, stop="\n")    lm += "\nAge: " + gen("age", regex=r"[0-9]+", max_tokens=3)    return lm # Use the functionlm = models.Transformers("microsoft/Phi-4-mini-instruct")lm = generate_person(lm) print(lm["name"])print(lm["age"])``` **Stateful Functions:** ```python@guidance(stateless=False)def react_agent(lm, question, tools, max_rounds=5):    """ReAct agent with tool use."""    lm += f"Question: {question}\n\n"     for i in range(max_rounds):        # Thought        lm += f"Thought {i+1}: " + gen("thought", stop="\n")         # Action        lm += "\nAction: " + select(list(tools.keys()), name="action")         # Execute tool        tool_result = tools[lm["action"]]()        lm += f"\nObservation: {tool_result}\n\n"         # Check if done        lm += "Done? " + select(["Yes", "No"], name="done")        if lm["done"] == "Yes":            break     # Final answer    lm += "\nFinal Answer: " + gen("answer", max_tokens=100)    return lm``` ## Backend Configuration ### OpenAI (remote — unconstrained only) > Remote API backends cannot do constrained generation (regex/select/grammar);> use them only for plain chat/`gen()`. For constraints, use a local backend. ```pythonfrom guidance import models lm = models.OpenAI(    model="gpt-4o-mini",    api_key="your-api-key"  # Or set OPENAI_API_KEY env var)``` ### Local Models (Transformers) ```pythonfrom guidance.models import Transformers lm = Transformers(    "microsoft/Phi-4-mini-instruct",    device="cuda"  # Or "cpu")``` ### Local Models (llama.cpp) ```pythonfrom guidance.models import LlamaCpp lm = LlamaCpp(    model_path="/path/to/model.gguf",    n_ctx=4096,    n_gpu_layers=35)``` ## Common Patterns ### Pattern 1: JSON Generation ```pythonfrom guidance import models, gen, system, user, assistant lm = models.Transformers("microsoft/Phi-4-mini-instruct") with system():    lm += "You generate valid JSON." with user():    lm += "Generate a user profile with name, age, and email." with assistant():    lm += """{    "name": """ + gen("name", regex=r'"[A-Za-z ]+"', max_tokens=30) + """,    "age": """ + gen("age", regex=r"[0-9]+", max_tokens=3) + """,    "email": """ + gen("email", regex=r'"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}"', max_tokens=50) + """}""" print(lm)  # Valid JSON guaranteed``` ### Pattern 2: Classification ```pythonfrom guidance import models, gen, select lm = models.Transformers("microsoft/Phi-4-mini-instruct") text = "This product is amazing! I love it." lm += f"Text: {text}\n"lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="sentiment")lm += "\nConfidence: " + gen("confidence", regex=r"[0-9]+", max_tokens=3) + "%" print(f"Sentiment: {lm['sentiment']}")print(f"Confidence: {lm['confidence']}%")``` ### Pattern 3: Multi-Step Reasoning ```pythonfrom guidance import models, gen, guidance @guidancedef chain_of_thought(lm, question):    """Generate answer with step-by-step reasoning."""    lm += f"Question: {question}\n\n"     # Generate multiple reasoning steps    for i in range(3):        lm += f"Step {i+1}: " + gen(f"step_{i+1}", stop="\n", max_tokens=100) + "\n"     # Final answer    lm += "\nTherefore, the answer is: " + gen("answer", max_tokens=50)     return lm lm = models.Transformers("microsoft/Phi-4-mini-instruct")lm = chain_of_thought(lm, "What is 15% of 200?") print(lm["answer"])``` ### Pattern 4: ReAct Agent ```pythonfrom guidance import models, gen, select, guidance @guidance(stateless=False)def react_agent(lm, question):    """ReAct agent with tool use."""    tools = {        "calculator": lambda expr: eval(expr),        "search": lambda query: f"Search results for: {query}",    }     lm += f"Question: {question}\n\n"     for round in range(5):        # Thought        lm += f"Thought: " + gen("thought", stop="\n") + "\n"         # Action selection        lm += "Action: " + select(["calculator", "search", "answer"], name="action")         if lm["action"] == "answer":            lm += "\nFinal Answer: " + gen("answer", max_tokens=100)            break         # Action input        lm += "\nAction Input: " + gen("action_input", stop="\n") + "\n"         # Execute tool        if lm["action"] in tools:            result = tools[lm["action"]](lm["action_input"])            lm += f"Observation: {result}\n\n"     return lm lm = models.Transformers("microsoft/Phi-4-mini-instruct")lm = react_agent(lm, "What is 25 * 4 + 10?")print(lm["answer"])``` ### Pattern 5: Data Extraction ```pythonfrom guidance import models, gen, guidance @guidancedef extract_entities(lm, text):    """Extract structured entities from text."""    lm += f"Text: {text}\n\n"     # Extract person    lm += "Person: " + gen("person", stop="\n", max_tokens=30) + "\n"     # Extract organization    lm += "Organization: " + gen("organization", stop="\n", max_tokens=30) + "\n"     # Extract date    lm += "Date: " + gen("date", regex=r"\d{4}-\d{2}-\d{2}", max_tokens=10) + "\n"     # Extract location    lm += "Location: " + gen("location", stop="\n", max_tokens=30) + "\n"     return lm text = "Tim Cook announced at Apple Park on 2024-09-15 in Cupertino." lm = models.Transformers("microsoft/Phi-4-mini-instruct")lm = extract_entities(lm, text) print(f"Person: {lm['person']}")print(f"Organization: {lm['organization']}")print(f"Date: {lm['date']}")print(f"Location: {lm['location']}")``` ## Best Practices ### 1. Use Regex for Format Validation ```python# ✅ Good: Regex ensures valid formatlm += "Email: " + gen("email", regex=r"[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}") # ❌ Bad: Free generation may produce invalid emailslm += "Email: " + gen("email", max_tokens=50)``` ### 2. Use select() for Fixed Categories ```python# ✅ Good: Guaranteed valid categorylm += "Status: " + select(["pending", "approved", "rejected"], name="status") # ❌ Bad: May generate typos or invalid valueslm += "Status: " + gen("status", max_tokens=20)``` ### 3. Leverage Token Healing ```python# Token healing is enabled by default# No special action needed - just concatenate naturallylm += "The capital is " + gen("capital")  # Automatic healing``` ### 4. Use stop Sequences ```python# ✅ Good: Stop at newline for single-line outputslm += "Name: " + gen("name", stop="\n") # ❌ Bad: May generate multiple lineslm += "Name: " + gen("name", max_tokens=50)``` ### 5. Create Reusable Functions ```python# ✅ Good: Reusable pattern@guidancedef generate_person(lm):    lm += "Name: " + gen("name", stop="\n")    lm += "\nAge: " + gen("age", regex=r"[0-9]+")    return lm # Use multiple timeslm = generate_person(lm)lm += "\n\n"lm = generate_person(lm)``` ### 6. Balance Constraints ```python# ✅ Good: Reasonable constraintslm += gen("name", regex=r"[A-Za-z ]+", max_tokens=30) # ❌ Too strict: May fail or be very slowlm += gen("name", regex=r"^(John|Jane)$", max_tokens=10)``` ## Comparison to Alternatives | Feature | Guidance | Instructor | Outlines | LMQL ||---------|----------|------------|----------|------|| Regex Constraints | ✅ Yes | ❌ No | ✅ Yes | ✅ Yes || Grammar Support | ✅ CFG | ❌ No | ✅ CFG | ✅ CFG || Pydantic Validation | ❌ No | ✅ Yes | ✅ Yes | ❌ No || Token Healing | ✅ Yes | ❌ No | ✅ Yes | ❌ No || Local Models | ✅ Yes | ⚠️ Limited | ✅ Yes | ✅ Yes || API Models | ✅ Yes | ✅ Yes | ⚠️ Limited | ✅ Yes || Pythonic Syntax | ✅ Yes | ✅ Yes | ✅ Yes | ❌ SQL-like || Learning Curve | Low | Low | Medium | High | **When to choose Guidance:**- Need regex/grammar constraints- Want token healing- Building complex workflows with control flow- Using local models (Transformers, llama.cpp)- Prefer Pythonic syntax **When to choose alternatives:**- Instructor: Need Pydantic validation with automatic retrying- Outlines: Need JSON schema validation- LMQL: Prefer declarative query syntax ## Performance Characteristics **Latency Reduction:**- 30-50% faster than traditional prompting for constrained outputs- Token healing reduces unnecessary regeneration- Grammar constraints prevent invalid token generation **Memory Usage:**- Minimal overhead vs unconstrained generation- Grammar compilation cached after first use- Efficient token filtering at inference time **Token Efficiency:**- Prevents wasted tokens on invalid outputs- No need for retry loops- Direct path to valid outputs ## Resources - **Documentation**: https://guidance.readthedocs.io- **GitHub**: https://github.com/guidance-ai/guidance (18k+ stars)- **Notebooks**: https://github.com/guidance-ai/guidance/tree/main/notebooks- **Discord**: Community support available ## See Also - `references/constraints.md` - Comprehensive regex and grammar patterns- `references/backends.md` - Backend-specific configuration- `references/examples.md` - Production-ready examples   
Discovery context

Discovered by repository scan. No exact path reference found in the snapshot’s root AGENTS.md.