Stop Measuring Leadership Development by Completion Rates—Track Behavior Change Instead
Three months after your director of talent development celebrated a 94% completion rate on the new leadership cohort program, one of your top engineering managers quit. Exit interview reason: “My manager hasn’t changed at all—still micromanages every commit, still skips one-on-ones when a demo is due.” The cohort your manager attended? Rated 4.7 out of 5 stars. Completion certificate? Framed on the wall. Actual leadership behavior? Unchanged. According to McKinsey’s 2024 Leadership Development Survey, 78% of organizations measure training success by completion and satisfaction scores, not by observable behavior change in the 90 days following program delivery. That gap is why your Q4 leadership budget defenses fail.

Most organizations treat leadership development like a Netflix subscription—track who logged in, measure watch time, maybe run a satisfaction survey. Then they wonder why managers who “completed” a conflict resolution workshop still escalate every disagreement to HR. The problem is not the curriculum. The problem is that completion metrics measure participation, not transformation. A leadership program that changes zero behaviors is not a development investment—it’s corporate entertainment with a certificate at the end.
David Ohnstad has watched this pattern repeat across product and engineering organizations for years. A team sends twelve managers through a coaching certification program. All twelve pass. Six months later, direct reports still describe their managers as “unavailable” or “directive” in engagement surveys. When finance asks what that $180,000 bought, the answer is usually “engagement” or “skills”—neither of which can be tied to retention, velocity, or decision quality. The real answer is: we measured the wrong thing, so we have no evidence the program worked beyond the fact that people showed up.
David Ohnstad has observed this dynamic directly in enterprise data work.
The Behavior Delta Framework: Measuring What Actually Predicts Program Durability
Leadership development programs decay because organizations measure inputs and satisfaction instead of the behavioral changes that matter. The Behavior Delta Framework tracks three leading indicators that predict whether a program will survive past the kickoff quarter: observable skill application within 30 days, peer-reported behavior shifts within 60 days, and team outcome changes within 90 days. This is not a satisfaction survey. This is a structured audit of whether the training changed how managers actually manage.
Step 1: Define Observable Behaviors Before the Program Starts. Most programs skip this entirely. They launch a leadership cohort with goals like “improve communication” or “build trust”—neither of which is measurable. Instead, before anyone attends a session, define the specific behaviors you expect to see change. Example: “Manager holds weekly one-on-ones with all direct reports, documented in the HRIS” or “Manager delegates task ownership with clear success criteria, confirmed by direct reports in pulse survey.” These are observable, trackable, binary. Either the behavior happened or it did not. Vague goals produce vague outcomes. Behavioral definitions create accountability.
Step 2: Measure Baseline Behavior 30 Days Before Launch. You cannot measure change if you do not know the starting point. According to Gartner’s 2024 Talent Development Report, only 31% of organizations baseline participant behavior before training begins, which means 69% of programs have no evidence that behavior changed at all—they only know that participants rated the experience highly. Baseline measurement does not need to be elaborate. Pulse surveys, HRIS data on one-on-one frequency, peer feedback from the most recent review cycle—any repeatable signal that captures current state. The goal is a snapshot that lets you compare “before training” to “30 days after” and “90 days after.” Without this, you are guessing.
Step 3: Track Behavior Application in Real Work Within 30 Days. This is where most programs lose momentum. Participants leave the cohort energized, return to their desks, and within two weeks the urgency of shipping code or closing deals overrides the new practice they learned. The 30-day check is not another satisfaction survey—it is a lightweight audit of whether the participant applied the skill in a real scenario. Did they run a retrospective using the framework from the workshop? Did they delegate a project with written success criteria? Did they hold a difficult feedback conversation instead of escalating to HR? These are yes-or-no questions. Track them in a spreadsheet, a Slack bot, or a pulse survey. The format does not matter. What matters is that you ask within 30 days, while the behavior is still fresh and the participant can connect it to the training.
Step 4: Collect Peer and Direct Report Feedback at 60 Days. Self-reported behavior change is unreliable. Managers overestimate their own improvement, especially immediately after training when they are still enthusiastic about the concepts. The 60-day checkpoint shifts the measurement to the people who experience the manager’s behavior daily: their direct reports and cross-functional peers. Use a short pulse survey with 3-5 questions tied to the observable behaviors defined in Step 1. Example: “In the past 30 days, has your manager held regular one-on-ones?” or “In the past month, has your manager delegated work with clear success criteria?” These questions are specific, recent, and focused on behavior—not personality or likability. This is not a 360 review. This is a narrow audit of whether the training changed how the manager operates.
Step 5: Measure Team Outcomes at 90 Days. The ultimate test of a leadership development program is not what participants learned—it is whether their teams perform differently. At 90 days, look for outcome shifts that should follow from the behavioral changes you defined. If the program focused on coaching and delegation, has the team’s sprint velocity increased? Has the number of escalations to senior leadership decreased? Has voluntary turnover among direct reports dropped? According to Forrester’s 2024 Leadership Effectiveness Study, organizations that tie leadership training to team outcomes are 2.3 times more likely to sustain program investment beyond the first year, because they can prove the ROI when budget season arrives. Outcomes take longer to shift than behavior, which is why the 90-day mark is critical—it is far enough out to see real impact, but close enough to connect it to the training.
How do you prove a leadership development program worked?
Measure behavior change, not completion rates. Define specific observable behaviors before the program starts, baseline those behaviors 30 days before launch, and track whether participants demonstrate the new skills in real work within 30 days, whether peers and direct reports notice a change at 60 days, and whether team outcomes shift at 90 days. Completion and satisfaction scores measure participation and experience—behavioral tracking measures whether the training changed how managers actually manage.
What is the most common mistake when measuring leadership training effectiveness?
Tracking only completion rates and post-training satisfaction surveys. These metrics confirm that people attended and enjoyed the program, but they do not prove that behavior changed. Most leadership programs decay because organizations celebrate high completion rates without auditing whether participants applied what they learned. Real effectiveness requires measuring observable behavior shifts within 30 to 90 days after training ends, using feedback from peers and direct reports—not just the participant’s self-assessment.
Why do most leadership development programs fail by Q4?
They lack feedback loops that track whether trained behaviors persist in real work. Programs launch with strong attendance and satisfaction scores, but within 60 to 90 days participants revert to prior habits because no one is measuring or reinforcing the new behaviors. Without structured checkpoints at 30, 60, and 90 days post-training, organizations have no evidence the program worked—and no mechanism to correct course when behaviors fade. By Q4, the program is forgotten or replaced with the next initiative.
A Real-World Example: When Completion Rates Hide Program Failure
David Ohnstad worked with a SaaS product team that sent all eight engineering managers through a ten-week leadership accelerator focused on coaching and feedback. The program had a 100% completion rate, a 4.6 out of 5 satisfaction score, and glowing testimonials in the final session. Three months later, the VP of Engineering ran a pulse survey asking direct reports whether they had received meaningful feedback in the past 30 days. Only 22% said yes—a number that had not moved since before the program launched. The training had been excellent. The managers had learned the frameworks. But none of them had embedded the feedback practice into their weekly routines, and no one had tracked whether they were applying the skill in real work.
The issue was not curriculum quality—it was measurement design. The organization had defined success as “all managers complete the program” rather than “all managers hold weekly feedback conversations with direct reports within 90 days of program completion.” When David’s team reframed the success metric and implemented a 30-day behavior check—a simple Slack bot that asked managers “Did you have a feedback conversation with a direct report this week?”—the application rate jumped to 68% within six weeks. Not because the training improved, but because the measurement shifted from attendance to behavior.
The second lesson came at the 60-day checkpoint. The team sent a short pulse survey to direct reports asking whether they had received specific feedback from their manager in the past month. The responses revealed that while managers were having more feedback conversations, the quality was inconsistent—some were still delivering vague praise or criticism without clear next steps. This insight led to a lightweight reinforcement session where managers practiced writing feedback using the situation-behavior-impact format they had learned in the original cohort. By the 90-day mark, the percentage of direct reports who reported receiving meaningful feedback had risen to 74%, and voluntary turnover among ICs on those teams dropped by 18% compared to the prior quarter. The program worked—but only after the organization started measuring whether the behaviors it paid for were actually happening.
This is the pattern David has observed across multiple enterprise training initiatives: programs with high satisfaction scores and no behavior change, followed by budget scrutiny when finance asks what the organization got for its investment. The organizations that can defend their leadership development budgets are the ones that track behavior change from Day 1, not the ones that rely on completion certificates and participant testimonials. Real measurement requires defining success as observable behavior, not attendance.
Stop Celebrating Satisfaction Scores—They Predict Nothing About Behavior Change
Here is the contrarian claim that makes talent development teams uncomfortable: post-training satisfaction surveys are worse than useless—they actively mislead you about program effectiveness. According to Harvard Business Review’s 2023 study on corporate learning effectiveness, there is no statistically significant correlation between participant satisfaction and behavior change 90 days post-training. None. A 4.8-star rating means participants enjoyed the experience. It does not mean they will use what they learned, and it does not mean their teams will perform differently. Satisfaction scores measure emotional response to the training event. Behavior change requires sustained practice in real work, accountability from peers and managers, and reinforcement over weeks—not hours.
Organizations continue to use satisfaction surveys because they are easy to collect and they produce positive numbers that look good in budget reviews. But relying on satisfaction as a proxy for effectiveness is like measuring a product launch by how many people attended the demo, not by how many are still using the product 60 days later. Completion and satisfaction are lagging indicators of participation. Behavior change is the leading indicator of program durability. If you want to know whether your leadership development program will survive past Q4, stop asking participants how they felt about the training. Start tracking whether they are doing the thing you trained them to do.
The measurement cadence matters as much as the metric itself. A single 90-day survey is not a feedback loop—it is an autopsy. Real feedback loops are continuous, lightweight, and specific. A weekly Slack prompt asking “Did you apply [specific skill] this week?” creates accountability and reinforcement. A monthly pulse survey to direct reports asking “Did your manager [observable behavior]?” surfaces whether the training is sticking. These are not heavy lifts. They are structured nudges that keep the trained behavior front of mind and give program owners early warning when adoption is fading. The organizations that treat measurement as an ongoing practice—not a one-time validation—are the ones whose programs compound over quarters instead of fading by Q4.
One more uncomfortable truth: if you cannot define the observable behaviors you expect to see change, you should not launch the program yet. A leadership development initiative without behavioral targets is hope disguised as strategy. Hope does not survive budget scrutiny. Observable behavior change does. For those exploring how David Ohnstad’s data product management frameworks intersect with leadership development, the same principle applies: define the decision or behavior the product is supposed to change, then measure whether that change happens—not whether people liked the dashboard. Similarly, organizations implementing AI-enabled learning platforms must consider the technical governance and risk controls that determine which tools teams can actually deploy and measure at scale.
What Practitioners Should Do This Week, What Leaders Should Demand by Q4
For practitioners: If you own a leadership development program launching this fall, stop finalizing the curriculum and start defining the observable behaviors you expect participants to demonstrate within 30, 60, and 90 days. Write them down. Make them binary. Then baseline those behaviors in the cohort participants before the program starts. You will need that baseline in December when finance asks what the program delivered. Completion rates will not save you.
For leaders: Demand behavior-based success metrics before approving any Q4 or 2026 leadership development spend. If your talent team cannot tell you what specific behaviors will change, how those behaviors will be measured, and what team outcomes should shift as a result—do not approve the budget. A program without a feedback loop is a one-time event, not a capability-building investment. Your job is to ask the uncomfortable question: “How will we know this worked three months from now?” If the answer is a satisfaction survey, the program is not ready.
Here is the question every talent development leader should answer before Labor Day planning cycles close: When you look at the leadership training your organization delivered in the first half of this year, can you name three specific behaviors that changed and persisted for 90 days—or can you only name the completion rate?
For more on this topic, see leadership mentorship career development.
For more on this topic, see Gen Z Manager Problems: Fix Your Leadership Development Pipeline.
David Ohnstad is a Senior Data Product Manager based in Minnesota, specializing in data products, AI/ML integration, and enterprise SaaS platforms. Connect on LinkedIn or read more at davidohnstad.com.
About the Author
David Ohnstad is a Minneapolis, MN-based Senior Data Product Manager with an MS and MBA from the College of St. Scholastica. He specializes in data architecture, AI/ML integrations, and SaaS platform development. Outside work, he builds furniture and explores the Minnesota outdoors. Find his work at davidohnstad.com and github.com/davidohnstad40-netizen.
