Ouroboros 🤖 Agent Open source
Spec-first workflow engine that runs an interview-evaluate-evolve loop over coding agents
- GitHub stars
- 6.2k
- Stars this week
- +69
- Forks
- 622
- Licence
- MIT
- Last push
- 2026-10-02
- Maintainer
- Q00
git clone https://github.com/Q00/ouroboros && cp ouroboros/*.md ~/.claude/agents/Third-party subagents & agents run with your permissions. Read the source before installing, and prefer pinned versions.
Works with
About Ouroboros
What it does
Ouroboros is a specification-first workflow engine that sits on top of your existing coding agent (Claude Code, Codex CLI, OpenCode, Gemini, Kiro, Copilot and others) rather than replacing it. Instead of letting an agent start coding from a vague prompt, it runs a structured loop: Interview (Socratic questioning to expose hidden assumptions), Seed (turning answers into an immutable specification), Execute (implementation via "Double Diamond" decomposition), Evaluate (mechanical, semantic and consensus checks), and Evolve (feeding results back into the next round).
What is inside
- A "Ralph" command that runs the full loop persistently until convergence, with event-sourced state and replay across sessions
- Quantified ambiguity scoring across goal, constraints, success criteria and context, with a numeric threshold before execution proceeds
- Ontology-convergence and drift measurements (weighted across goal, constraints and schema similarity) to detect when a spec has stabilized or is oscillating
- A "PAL Router" that auto-escalates to more expensive models on failure and downgrades again on success
- An orchestrator layer abstracting over more than eight different coding-agent runtimes
Works with
Documented to run on top of Claude Code, Codex CLI, OpenCode, Gemini, Kiro, GitHub Copilot and several other agent CLIs as pluggable backends.
How to install or connect
A one-command shell/PowerShell setup is documented; it requires Python 3.12+.
Maintenance and safety
MIT licensed, with over 2,200 commits and roughly 6,000 stars — an active project with its architecture openly documented, including specific module responsibilities. Its "self-improving" language is worth reading carefully: the loop refines the specification and drives re-execution, it does not mean the underlying LLM itself learns or changes weights, so treat the marketing framing with a bit of skepticism even though the mechanism itself is real and documented.
Who should use it
Teams who want more rigor than "prompt and hope" before letting a coding agent build a nontrivial feature, and who are willing to sit through an upfront interview/spec phase in exchange for more convergent, checked output.
Pros
- Concrete, documented mechanism (ambiguity scoring, ontology convergence, drift measurement) behind its evolve loop rather than vague hype
- Works across 8+ existing coding-agent CLIs instead of locking you into one
- Active development with full architecture docs and event-sourced replay
Cons
- "Self-improving" framing overstates what actually happens (spec refinement, not model learning)
- Upfront interview/spec process adds overhead that's unnecessary for small, simple tasks
Similar subagents & agents
All agent workflows & frameworks →Spec Kit 🤖 AgentFree
GitHub's toolkit for spec-driven development with AI coding agents
Task Master 🤖 AgentFree
AI task management that turns a PRD into tasks your coding agent works through
Fast Agent 🤖 AgentOpen source
Python framework for building, orchestrating and evaluating MCP-native AI agents
OpenSpec 🤖 AgentFree
Lightweight spec-driven development: agree on changes before the agent codes
CC Safety Net 🤖 AgentOpen source
Pre-execution guard that blocks destructive commands and secret access for coding agents
BMAD Method 🤖 AgentFree
Agile AI-driven development with analyst, PM, architect, developer and UX agents