Agent self-reflection is an agent design pattern where an AI agent critiques its own intermediate reasoning, tool-use trajectory, or final answer, then writes a revised plan or output based on that critique.
What is Agent Self-Reflection?
Self-reflection adds an explicit review step inside the agent loop: draft a plan/answer, critique it using a rubric (constraints, evidence, safety), then revise and optionally re-run tools. It can be implemented with the same model under a critic prompt or with a separate judge model. In high-stakes flows, reflection can trigger human review.
Where it’s used and why it matters
Self-reflection is used in research agents, coding agents, and enterprise automation where errors compound across steps. It can improve success rate and reduce unsafe tool actions by acting as a quality gate, but it increases cost and latency.
Examples of Agent Self-Reflection in Practice
- RAG assistant: verify claims are supported; retrieve more if not.
- Coding agent: run tests; revise based on failures.
- Ops agent: check permissions/reversibility before approval.
FAQs
How is this different from chain-of-thought? Chain-of-thought is reasoning within one completion; reflection critiques and revises a prior draft/trajectory.
What should the rubric include? Constraints, evidence, tool safety, budgets, and stop criteria.
Can it reduce hallucinations? Yes, when reflection checks grounding and forces abstention or more retrieval.
How can I practice? Add a critic pass to a tool-using agent and measure success rate vs. added latency.