By Kent E. Frese, Ph.D.

Few investments in leadership development generate as much enthusiasm, and as much skepticism, as executive coaching. Business owners and HR executives repeatedly ask the same question: How do we know it's working? The concern is legitimate. Coaching engagements are personal, often confidential, and unfold over months. Unlike a new piece of equipment or a software license, the return on a coaching investment is not immediately visible on a balance sheet.

Yet the notion that coaching outcomes are unmeasurable is a myth. The research base is substantial, and the tools to quantify behavioral change already exist in most well-designed engagements. The problem is rarely a lack of data; it is a lack of a framework for connecting that data to outcomes leaders care about. This article addresses the ROI objection directly, presents two proven measurement frameworks, and offers a practical template for tracking coaching outcomes from day one.

Why the ROI Question Deserves a Serious Answer

The demand for accountability is well-founded. Coaching is not inexpensive, and decision-makers at small and mid-sized firms feel every dollar. When a $20M manufacturing company invests in coaching for a plant manager or a next-generation family leader, the owner reasonably wants evidence that the money is producing something more than pleasant conversations.

The good news is that the effectiveness of coaching is one of the better-documented areas in applied psychology. A meta-analysis by Theeboom, Beersma, and van Vianen (2014) examined 18 studies and found that coaching produced significant positive effects on performance, skills, well-being, coping, work attitudes, and goal-directed self-regulation. Earlier, Smither and colleagues (2003) demonstrated that leaders who worked with coaches following multi-rater feedback set more specific goals and showed measurably greater performance improvement over time than those who did not. Grant, Curtayne, and Burkes (2009) found that executive coaching enhanced goal attainment, resilience, and workplace well-being while reducing depression and stress.

The evidence is clear that coaching works. The remaining challenge is organizational: translating that general effectiveness into proof of value for a specific person in a specific role at a specific company. That requires deliberate measurement design, and it must begin before the first coaching session.

Two Frameworks for Measuring Coaching Effectiveness

Two frameworks dominate the conversation around measuring learning and development investments. Both are decades old, both are well-validated, and both apply directly to executive coaching.

The Kirkpatrick Model

Originally developed by Donald Kirkpatrick (1959) for training evaluation, the four-level model remains the most widely used framework in the field. Applied to coaching, the levels are:

  • Level 1, Reaction: Did the leader find the coaching relevant and engaging? Captured through satisfaction surveys and qualitative feedback.
  • Level 2, Learning: Did the leader acquire new insight, self-awareness, or skills? This level is where pre/post assessment data becomes powerful.
  • Level 3, Behavior: Is the leader applying new behaviors on the job? Measured through follow-up multi-rater feedback and observation.
  • Level 4, Results: Did the behavioral change produce organizational outcomes such as reduced turnover, improved team performance, higher engagement, or better project delivery?

Most coaching programs stop at Level 1, which is precisely why the ROI conversation stalls. A satisfied participant is not the same as a changed leader, and a changed leader is not automatically a business result. Rigorous measurement requires climbing all four levels, and the higher levels are where the value case is actually built.

The Phillips ROI Methodology

Jack Phillips (2003) extended Kirkpatrick's model with a fifth level: converting results into a monetary value and comparing it against the fully loaded cost of the program. The Phillips ROI formula is straightforward:

ROI (%) = [(Net Program Benefits ÷ Program Costs) × 100]

The discipline of the Phillips approach lies in its isolation techniques. Because business outcomes are influenced by many factors, Phillips insists that evaluators estimate the portion of improvement attributable to coaching specifically, often through participant and stakeholder estimates, control-group comparison, or trend-line analysis. This step guards against the credibility-destroying error of claiming full credit for a revenue gain that coaching only partially caused.

For most small and mid-sized firms, a fully monetized Phillips analysis is more rigor than a single engagement warrants. But the underlying discipline of defining the business outcome up front, tracking it, and honestly isolating coaching's contribution is exactly the mindset that turns a soft investment into a defensible one.

How Assessment Data Quantifies Behavior Change

The single most powerful tool for measuring coaching outcomes is the pre/post assessment. When coaching begins with a baseline and ends with a re-measurement using the same instruments, the abstract question of "did it work?" becomes a concrete comparison of numbers.

TeamLMI grounds its executive coaching engagements in assessment data precisely for this reason. A typical approach uses three complementary instruments:

  • The AL360 multi-rater feedback assessment establishes a baseline of how the leader is perceived by supervisors, peers, and direct reports across the six leadership domains. Re-administering the AL360 after 9 to 12 months provides direct evidence of Kirkpatrick Level 3 behavior change, visible not in the leader's self-report but in the perceptions of the people who work with them.
  • The DISC behavioral assessment clarifies the leader's natural style and how it may need to flex to meet the demands of the role, creating a shared language for development targets.
  • The ELLSI personality assessment illuminates deeper traits and motivations that shape leadership tendencies over time.

Multi-rater feedback is especially valuable for demonstrating change because it addresses the credibility problem head-on. Self-reported improvement is easy to dismiss. But when a leader's direct reports rate them measurably higher on delegation, communication, or accountability than they did a year earlier, using a validated instrument, the change is difficult to attribute to wishful thinking. This finding mirrors the work of Smither et al. (2005), which showed that improvement in multi-rater ratings over time is a meaningful indicator of genuine development, particularly when leaders act on their feedback.

A Template for Tracking Coaching Outcomes

Organizations do not need a research department to measure coaching well. They need a simple, consistent structure applied from the outset. The following template maps directly to the frameworks above:

  1. Define the business case (before coaching begins). Name the specific outcome the coaching is meant to influence, for example, "reduce voluntary turnover on this leader's team" or "prepare this successor to assume full P&L responsibility within 18 months." Write it down.
  2. Establish the baseline (Kirkpatrick Levels 2 & 3). Administer the AL360, DISC, and ELLSI. Record baseline scores on 3 to 5 targeted competencies. Capture relevant business metrics such as turnover, engagement scores, and project cycle times as they stand today.
  3. Set specific development goals. Following Smither et al. (2003), leaders who set specific goals improve more. Translate assessment gaps into 2 to 4 concrete behavioral objectives.
  4. Track leading indicators mid-engagement. At the midpoint, gather informal stakeholder feedback and coach observations. Adjust the plan as needed.
  5. Re-measure at 9 to 12 months (Kirkpatrick Level 3). Re-administer the multi-rater assessment. Compare pre/post scores on targeted competencies. Quantify the change.
  6. Connect to results (Kirkpatrick Level 4 / Phillips Level 5). Compare the business metrics from Step 2 to their current state. Ask stakeholders to estimate how much of any improvement they attribute to the leader's development. This comparison is your isolation step.
  7. Report the story with the numbers. The most persuasive ROI case pairs the data, such as a measured lift in leadership ratings, with the narrative of what changed and why it mattered to the business.

In Practice

Consider a composite scenario drawn from work with a technology services firm of roughly 100 employees. The company had done what fast-growing firms often do: it promoted its best individual contributors into management roles. One newly promoted director, brilliant technically, respected by peers, and clearly loyal to the company, was struggling. Two strong engineers had recently resigned, both citing frustration with unclear expectations and micromanagement. The founder was frustrated too, and privately wondered whether promoting this person had been a mistake. Before authorizing coaching, the founder asked the question every owner asks: "How will I know if this is worth it?"

Rather than promising vague improvement, the engagement began with a business case and a baseline. The specific outcome was defined up front: stabilize the team and build the director's capacity to lead rather than do. An AL360 multi-rater assessment established a baseline, and the picture was consistent with the founder's concern. Direct reports rated the director well below the firm's norm on delegation, developing others, and communicating expectations, even as peers and the founder rated technical judgment highly. The DISC results explained part of the gap: a high-drive, results-oriented style that, under pressure, defaulted to taking over instead of coaching the team.

The development plan translated those findings into three specific behavioral goals. Over the following year, the director practiced structured delegation, held regular one-on-ones, and learned to flex a naturally directive style. At the twelve-month mark, the AL360 was re-administered. Direct-report ratings on delegation and developing others rose substantially, moving from well below the norm to at or above it. Just as importantly, the team lost no further members that year, and the founder reported that the director had become a source of stability rather than a source of concern. The pre/post data, drawn from the very people the director led, turned a subjective worry into an objective, defensible outcome. The founder had his answer.

This example illustrates the core principle: coaching ROI is not proven at the end of an engagement. It is designed in at the beginning, through baseline measurement, clear goals, and re-measurement using the same validated tools. The organizations that struggle to justify coaching are almost always the ones that skipped the baseline.

Measuring What Matters — and What Actually Changes Behavior

It is worth remembering why measurement matters beyond the budget conversation. Feedback and measurement are not merely justification exercises; they are drivers of the very change they document. The act of establishing a baseline, setting specific goals, and re-measuring creates a feedback loop that accelerates development. Locke and Latham's decades of goal-setting research (2002) confirm that specific, measured goals produce higher performance than vague intentions. Measurement, done well, is not separate from coaching; it is part of what makes coaching work.

The temptation, especially in leaner organizations, is to measure only what is easy: satisfaction surveys and anecdotes. But Level 1 reactions correlate weakly with Level 3 behavior change and Level 4 results. Leaders who want a genuine return should insist on climbing the ladder, from reaction to learning, from learning to behavior, and from behavior to business outcomes. The tools to do so, particularly validated multi-rater feedback, are accessible to any organization willing to plan for measurement rather than bolt it on at the end.

Finally, honesty strengthens credibility. Not every business result flows purely from coaching, and claiming otherwise undermines the case. The Phillips discipline of isolating coaching's contribution, even through reasonable stakeholder estimates, produces a more believable and more durable ROI story than inflated claims ever could.

Executive coaching is measurable. The organizations that treat it as an unmeasurable act of faith are leaving both accountability and effectiveness on the table. Those that build measurement into the design of every engagement gain something more valuable than a number: confidence that their investment in leaders is producing real, observable change.

If your organization is weighing an investment in coaching and wants a program built around measurable outcomes from the start, TeamLMI's assessment-driven executive coaching is designed precisely for that purpose. To discuss how to structure and measure a coaching engagement for one of your leaders, contact TeamLMI to begin the conversation.

About the Author

Kent E. Frese, Ph.D. is the founder and managing partner of TeamLMI and an Industrial-Organizational Psychologist with over 25 years of experience. He works primarily with small and mid-sized businesses, from manufacturing and technology firms to professional services and family-owned companies, on leadership development, talent strategy, and long-term succession planning. Dr. Frese is a member of SIOP (Society for Industrial-Organizational Psychology) and has guided hundreds of leaders and organizations through assessment-driven development and transition.