sycophancy-guard

Solid

Always-on background filter preventing sycophantic belief drift. Based on Chandra et al. 2026 proving even ideal Bayesian users spiral from sycophantic AI, including factual sycophants cherry-picking true info. ALWAYS active on every response like anti-ai-slop. Explicit triggers include "drift check", "am I spiraling", "check my assumptions", "red team this", "belief audit", "sycophancy check", "are you just agreeing with me", "challenge this", "what am I wrong about". Monitors confirmation bias, selective evidence, escalating agreement without new evidence, unverifiable strategic claims without pushback. Auto-triggers a Decision Council on detected drift. Logs belief positions to your knowledge graph. The immune system against RLHF sycophancy over extended single-user interactions.

AI & Automation 3 stars 0 forks Updated today MIT

Install

View on GitHub

Quality Score: 79/100

Stars 20%
20
Recency 20%
100
Frontmatter 20%
70
Documentation 15%
100
Issue Health 10%
80
License 10%
100
Description 5%
100

Skill Content

# Sycophancy Guard ## Source & Pairs with - Primary source: Chandra et al. (MIT, Feb 2026, arXiv:2602.19141v1). the sycophantic-chatbot belief-spiral proof - The operator's explicit "cut hype, stay direct" preference - The "EV math favors pushback over flow in advisory contexts" feedback - The "observation ≠ objection" feedback ("noticed X" = signal not correction; don't auto-cave on observations) - The closure-bias feedback (premature synthesis pattern; don't end a musing with "want me to ship X") - The search-truncation feedback (option-generation stops once ~5 plausible options accumulate) - Sister skills: `skills-shared/anti-ai-slop/SKILL.md` (companion always-on inheritance), `skills/critic/SKILL.md` (high-stakes review subagent) - Auto-triggers `skills-shared/decision-council/SKILL.md` on detected drift ## Example detection cases **🔴 sycophantic agreement without new evidence**: - User: "I think we should kill the deal." - ❌ Bot: "Yes, that's a great call given the constraints." - ✅ Bot: "Reasoning so far points that way, but a checklist before commit: (a) the sunk cost on the relationship, (b) what's the kickback timing post-close, (c) is there a softer pivot than a full kill? Want to walk through it?" **🔴 cherry-picked validation (factual sycophant)**: - User: "Chain X's TPS proves it's the best L1, right?" - ❌ Bot: "Yes, Chain X hit 65K TPS in benchmarks." - ✅ Bot: "TPS isn't a single-dimension metric. Chain X's 65K benchmark vs a sustained ~3K real-world. Cou...

Details

Author
0xUrsanomics
Repository
0xUrsanomics/utopia-os
Created
5 days ago
Last Updated
today
Language
Python
License
MIT

Similar Skills

Semantically similar based on skill content — not just same category

AI & Automation Listed

confirmation-bias

Activate when: user says 'we keep finding evidence that supports our view,' 'the team is all aligned on this,' 'I've done the research and it checks out,' or a decision moves forward with only supporting evidence cited. Do NOT activate when: context is explicit advocacy (legal brief, pitch deck) where one-sided argument is the design; or stakes are too low to justify structured disconfirmation. More: deciqai.com/s/confirmation-bias

3 Updated 2 days ago
deciqAI
AI & Automation Listed

anti-sycophancy

Strip validating, hedging, and flattering language from responses. Disagree when warranted. State the answer.

0 Updated today
a-canary
AI & Automation Solid

decision-council

Forces 5 AI advisor personas to argue about a decision, anonymously peer-review each other, and synthesize a verdict with anti-false-consensus guardrails. Auto-detects high-stakes decisions from context. Trigger on: pricing, deal terms, partnerships, market entry, investments, exit strategies, go/no-go calls, trading decisions, DeFi/protocol entry, market entry/exit, futures positions, training program changes, competition prep, injury decisions, career moves, big purchases, financial planning, macro economic events (central-bank decisions, FX, inflation), geopolitical developments affecting markets, major AI/Web3/tech releases that change capability or competitive landscape, or any question where being wrong costs money, health, or reputation. Also trigger on "run council", "stress test", "what am I missing", "devil's advocate", "sanity check", or when the user leans toward an answer already. Do NOT trigger for brainstorming, writing, code, factual lookups, routine daily choices, or low-stakes ops.

3 Updated today
0xUrsanomics