Watch the full webinar on demand
Catch the full session with Megan Torrance, plus Q&A highlights.
SAVE NOW: New season, new skills. Earn the credential, with flexible payment options.
Blog
Share This Post
:format(webp))
:format(webp))
Insights from a recent Digital Learning Institute (DLI) webinar with Megan Torrance.
More than any other function in the business, learning and development seems to spend an outsized amount of energy justifying why it exists. Nobody in Accounts Payable has to prove the ROI of paying the bills every single day. Nobody in Facilities has to defend keeping the lights on. But L&D? L&D is expected to make its case, constantly, and in the age of Gen AI, that pressure has only intensified.
That's the opening provocation from our recent webinar with Megan Torrance on data, AI, and proving the value of learning. CEO and founder of Torrance Learning, and author of Agile for Instructional Designers, Data and Analytics for Instructional Designers, Making Sense of xAPI, and most recently The AI Implementation Guide for L&D, Megan Torrance brought nearly three decades of experience to a simple but uncomfortable question: if learning teams have more data than ever, why does it still fail to earn them credibility in the business?
Catch the full session with Megan Torrance, plus Q&A highlights.
Torrance opened with a real example: a client who had, by their own account, done everything right. They had data. They'd hired an analyst. That analyst was producing genuinely beautiful dashboards. And yet the learning team wasn't getting any more credibility, or any more answers, than before.
Look closely at a typical L&D dashboard and the pattern becomes obvious: page views, completions, who did what, on which platform, in which division. It's not wrong data, and it's not wrong to collect or report on it. It's just not enough. It tells you how much was consumed and by whom. It tells you almost nothing about whether any of it made a difference.
Torrance's framework for why this keeps happening comes down to four distinct failure points: the wrong data, data that isn't credible, data that isn't analyzed with any real rigor, and, in some cases, the wrong target altogether.
The instinct to reach for completions and satisfaction scores isn't random, it's what the field has always measured. Torrance cited her own 2003 research with the Learning Guild: when asked what they measure, most respondents pointed straight to completions and perceptions. Effectiveness and business impact trailed a distant third and fourth.
The uncomfortable quote she returned to, borrowed from Will Thalheimer's The CEO's Guide to Training, eLearning, and Work: evaluating training by how satisfied learners felt, how much they liked their instructor, and whether they'd recommend it, may be "the root of all evil in the training field." It rewards building training people enjoy, rather than training that actually builds competence.
The fix starts before a single slide gets built. Torrance's teams ask, as question one in every kickoff, "How are we going to measure success?" deliberately before design begins, not, as most L&D teams following the ADDIE model tend to do, as an afterthought squeezed into evaluation at the end. If a program is already live and that question was never asked, the fallback is to find whoever would notice if the program disappeared, and ask them what they'd expect to see, and how they'd know it was working.
Crucially, that means measuring people on their jobs, not on their learning. As Torrance put it, people don't show up to work to learn, they show up to do their jobs; learning is simply how they get there.
Even with the right target in view, data can fail on credibility. Torrance frames this through five lenses, the Five Vs of data:
Validity — is this actually a measure of what you think it's measuring? (Torrance's own go-to example: she holds a certification as a spray tan applicator, and has never in her life given, held, or received a spray tan. The certificate proves nothing about the skill.)
Volume — do you have enough data to draw a meaningful inference, or are you generalizing from 18 responses?
Variety — do you have enough different signals pointing the same direction to trust the conclusion?
Velocity — are you getting the data fast enough to act on it, or waiting a full quarter for numbers that are already stale?
Value — is this actually meaningful, or just measurable?
A recurring credibility killer is human data entry, both in the literal sense (manual entry is a known source of error) and in the psychological sense: people are not reliable narrators of their own behavior. Post-session survey scores cluster suspiciously at 4s and 5s. Self-reported "yes, I applied what I learned" answers reflect what people want to be seen doing, not necessarily what they did. Wherever possible, Torrance recommends pulling from systems of record instead of self-report, and being explicit with stakeholders about which is which.
Torrance draws a sharp line between two things L&D teams often blur together: reporting and analytics. Reporting is the steady drumbeat, how many people took this course this month, tracked because it stays interesting over time. Analytics is exploratory and episodic: does this program actually improve something, or not? A pharmaceutical company reports monthly sales figures every month, it doesn't re-run the clinical study on whether the drug works every month. L&D teams, Torrance argued, often forget that second category has its own separate rhythm and its own separate rigor.
Two analytical approaches came up as genuinely useful starting points:
Statistical process control, borrowed from operations, helps distinguish normal variation from a real signal. It tells you when a metric moving is worth getting excited about, and when it's just noise inside expected bounds.
Controlled comparison (A/B-style testing), which forces a discipline around statistical significance and confidence intervals, rather than eyeballing two numbers and declaring a training pilot a success.
This is also where Gen AI enters the picture, and where Torrance was most direct: "Your AI tools cannot tell you what is causal." AI tools are genuinely useful for spotting correlation, flagging statistical significance, and processing volumes of data no human team could get through manually. But they will confidently narrate cause and effect that isn't there, because, as Torrance put it, causal inference is itself a kind of hallucination, one AI tools are prone to and humans are too. Her practical rule: run correlation and significance questions through AI freely, but never let it make the causal call, and never rely on AI alone when the stakes involve real decisions about people's jobs. For anything high-stakes or genuinely complex, the right move is finding your organization's actual data scientists, not prompting harder.
One more practical habit worth stealing: don't automatically discard outliers. Splitting a group into quintiles, and paying attention to both the top and bottom 20%, often surfaces the most useful signal in the whole dataset. An outlier might be noise. It might also be the most interesting thing in the room.
L&D's fixation on ROI, Torrance argued, is largely inherited rather than chosen. The field was told to speak the language of the business, decided business means money, and concluded ROI must be the metric that matters most. But isolating the ROI of a single training intervention means isolating one course from Gilbert's behavioral engineering model, which identifies training and knowledge as just one of six factors driving job performance, and arguably the one L&D has the least leverage over compared to environment, incentives, tools, and feedback.
Real work is messy and entangled with other functions. Chase a perfectly clean, singularly attributable ROI number and you'll either fail to find it or shrink your scope so far you lose sight of actual performance. The alternative Torrance proposes isn't abandoning outcomes, it's building a reasonable, evidence-backed causal chain: decision competence and task competence in the training environment, moving toward on-the-job transfer, moving toward business effects, each link measured where it's actually measurable, rather than demanding one number that proves everything at once.
A question from an attendee mid-session captured a challenge nearly every instructional designer runs into: what do you do when you're only asked to build an e-learning course, but the real success metric, fewer calls to an admin center, lives entirely outside your control?
Torrance's answer: separate what you can instrument from what you can only request. Decision and task competence can be built and measured directly inside the learning experience itself, through scenario-based design, branching decisions, or AI-driven practice simulations that assess how someone actually responds to a situation, not just whether they picked the right multiple-choice answer. On-the-job transfer and downstream business effects require a different move entirely: a conversation with the business, asking them to track their own side of the causal chain, since that data usually lives in systems L&D doesn't own.
The point isn't that L&D can measure everything. It's that being disciplined and honest about the boundary of what you can measure, and partnering deliberately on the rest, does more for your credibility than quietly pretending completions cover it.
None of this is new information dressed up. What Torrance offered was a diagnostic: if your L&D data isn't landing with leadership, it's very likely failing at one of four specific, identifiable points, not because data doesn't matter, but because the wrong data, unreliable data, under-analyzed data, or data chasing the wrong target will never earn trust, no matter how polished the dashboard looks.
The path forward starts before a project kicks off, not after it launches: ask how success will be measured first, be honest about what's credible and what isn't, apply real analytical rigor (with AI as an assistant, never the final word on causation), and resist the pull toward a single, tidy ROI number when a causal chain tells a more honest story. Do what you can with what you have, where you are, and let that discipline be the thing that earns you the next, bigger dataset.
What are the four reasons L&D data fails to build credibility? Measuring the wrong data, data that isn't credible, data that isn't analyzed with sufficient rigor, and, in some cases, chasing the wrong overall target (like an isolated ROI figure).
What are the Five Vs of credible data? Validity (does it measure what you think it measures), Volume (enough data for a meaningful inference), Variety (multiple signals pointing the same way), Velocity (fast enough to act on), and Value (meaningful, not just measurable).
What's the difference between reporting and analytics? Reporting is the steady, recurring drumbeat of metrics that stay interesting over time, like monthly completions. Analytics is exploratory and episodic, answering a specific question like "does this program actually work," which doesn't need re-answering every month.
Can AI tools tell you what's causing a change in your data? No. AI tools can help identify correlation and statistical significance, but they cannot reliably determine causation, and will often confidently suggest causal relationships that aren't real. Treat causal claims from AI as a hypothesis to verify, not a conclusion.
Why shouldn't L&D chase ROI as the primary metric? Training is only one of several factors (alongside environment, incentives, tools, and feedback) that influence job performance, per Gilbert's behavioral engineering model. Isolating training's ROI from everything else influencing performance is often impractical, and pursuing it too narrowly can mean losing sight of the fuller picture. A defensible causal chain, from training to transfer to business effect, is often more honest and more useful.
What if I'm only asked to build an e-learning course, but the real success metric happens on the job? Focus on what you can directly instrument, decision and task competence within the learning experience itself, and treat on-the-job transfer and business outcomes as a partnership question for the business to help track, since that data typically lives in systems L&D doesn't control.
When should I stop trusting a self-reported survey result? Be skeptical whenever data relies on human recall or self-report, especially where people have an incentive to look good (e.g., "yes, I applied what I learned"). Where possible, pull equivalent data from a system of record instead.
Browse DLI's free resource library including expert-led masterclasses, webinars, ebooks, and toolkits to help you stay at the forefront of digital learning.