Methodology
Diagnose. Intervene. Measure. Hand off.
A delivery framework grounded in organizational psychology. Four phases. Each phase targets the human mechanisms that predict whether a team ships or stalls.
Phase 1
Diagnose
Before fixing anything, measure what's actually happening. Most delivery problems are misdiagnosed as process failures when they're human performance failures.
Psychological safety audit
Do people admit mistakes, or do they hide them? Edmondson (1999) found that teams with higher psychological safety report more errors but have fewer adverse events. If retros are polite and postmortems never name names, the team is hiding what it knows.
Goal specificity score
Locke and Latham (1990) demonstrated that specific, difficult goals with feedback outperform vague objectives. If your sprint goal starts with "Complete" and ends with a number, it's a container, not a goal. Score every active goal on specificity, difficulty, and feedback access.
Job characteristics check
Hackman and Oldham (1976) identified five dimensions that predict meaningful work: skill variety, task identity, task significance, autonomy, and feedback. Feature factories strip all five. Score the team's work against each dimension.
Team stage mapping
Tuckman (1965) described the stages every team moves through: forming, storming, norming, performing. Reorgs and new hires reset teams to forming. If your roadmap treats a storming team like a performing one, the timeline is fiction.
Phase 2
Intervene
The diagnosis names the mechanism. The intervention targets it directly. Each intervention is designed around a specific I/O psychology finding, applied at the right point in the delivery cycle.
Commitment design
Festinger (1957) showed that public commitments create psychological pressure to follow through, while private agreements carry no weight. Replace one-on-one alignment with public commitment forums where stakeholders state tradeoffs in front of each other. The conflict already exists. Better to have it in the room once than asynchronously across six weeks.
Goal restructuring
Goals improve performance through four mechanisms: direction, effort, persistence, and strategy (Locke & Latham). Vague goals collapse all four. Rewrite sprint goals as outcomes with observable completion criteria. If you can't tell whether the sprint succeeded from the goal alone, the goal is broken.
Decision hygiene
Janis (1982) identified eight symptoms of groupthink, from the illusion of invulnerability to self-censorship. Cohesive teams are most at risk. Install decision hygiene: written decision logs, assigned dissent roles during prioritization, anonymous pre-read feedback before roadmap commits.
Ownership calibration
Latané, Williams, and Harkins (1979) demonstrated that individuals reduce effort when responsibility is diffused across a group. Social loafing is not a character flaw. It's a structural response to ambiguous ownership. Every outcome needs one name. If the answer to "who owns this" is "the team," you have a diffusion problem.
Phase 3
Measure
Track the signals, not the ceremonies. Most delivery metrics measure activity. The research points to different indicators.
Goal specificity trend
Are sprint goals getting more specific over time? Score each sprint goal against the Locke & Latham criteria. A flat or declining trend means the intervention didn't stick.
Retro candor index
Are retros surfacing new categories of problems, or repeating the same safe topics? Edmondson's research predicts that candor increases as psychological safety grows. Track topic diversity and novelty.
Decision latency
How long from "we should decide this" to "we decided"? Long latencies signal a commitment problem: either ownership is ambiguous or stakeholders are avoiding public tradeoffs for fear of dissonance.
Cycle time by ownership clarity
Compare cycle time for items with clear owners versus items assigned to "the team." The gap tells you how much social loafing is costing you. If there's no gap, ownership is unambiguous and the diffusion intervention worked.
Phase 4
Hand off
The goal is not to make the team dependent on a consultant. It's to leave them with the diagnostic tools to self-correct. The framework belongs to the team.
Documented baselines
Every diagnosis produces a baseline: the psychological safety score, the goal specificity trend, the ownership map. The team can rerun the diagnostics themselves when something feels off.
Decision logs
Written, public, searchable records of who agreed to what and when. Cognitive dissonance doesn't require a facilitator. It requires a witness. The decision log is the witness.
Trigger thresholds
Specific, measurable conditions that signal a re-diagnosis is needed: retro candor drops below baseline for two consecutive sprints. Cycle time variance exceeds 40%. Goal specificity score drops. The team knows when to act.
Where agents fit
Pattern detection at scale.
AI agents handle the layer humans are bad at: sustained, unglamorous pattern detection across time and teams.
Agents can score sprint goals against specificity criteria automatically. They can detect declining retro candor across 12 sprints without getting bored. They can flag that a decision had no dissenting voice in the log. They can surface ownership ambiguity before the cycle time data makes it obvious.
What agents cannot do: experience cognitive dissonance when they're wrong. Feel the weight of a public commitment. Build psychological safety through vulnerability. Every agent-assisted signal needs a human owner who experiences the consequences of acting on it, or ignoring it.