A conversation with Diane Talley, Psychometrician at Meta
Bob Mosher (BM): Diane, please give us a little bit of your background, how you got here, and the role you play.
Diane Talley (DT): I have a PhD in educational measurement. In this field we focus on the science of measurement, including survey research methods. My specific expertise is really in competency testing. I spent around 25 years in certification and licensure, give or take some time in grad school.
I came to Meta four years ago to evaluate their approach to competency assessments aligned with training and make recommendations for improvement. Eight months later I became a full-time employee. I spent the last three and a half years building out strategy and implementation of an assessment ecosystem and improving both the way we do surveys and when we do surveys. I’ve also expanded into other areas of measurement such as KPIs/ROI.
BM: Brilliant.
DT: It’s been an exciting road. We’re doing a lot of work to build a system for doing virtual performance-based testing, which is much better than doing multiple-choice assessments. Within that system, we will use AI in a number of ways including question and scenario generation. Meta is very big on “dogfooding,” or using our own products, so AI is a part of everyday life.
BM: I know that strategy well. We did the same at Microsoft. We kicked the tires before we inflicted products on everybody else and it was always quite a journey.
Being an L&D professional for 43 years, I deeply appreciate who you are and the role you play. I took test writing courses in my undergraduate work and was a public-school teacher at the start of my career. Test writing was required, but it was multiple choice questions vs. the art of assessment, or should I say measurement, which is a better word.
As I got into corporate training, I realized test writing was a severely underused muscle. ROI came along pretty early into my career, and I just remember my colleagues and I saying, “What? We’re designers.” Or— I’m about to say something that will date me— “I used to be a secretary and now I’m writing Word courses, and I have to prove what?”
Unfortunately, I would argue that in the L&D industry, measurement is still a pretty weak area for many organizations. Super weak, right?
DT: I had no idea what I would find when I first started. I had previously focused on certification and licensure where training is separated from testing.
So, when I came to Meta, it was a little bit of a shock to me to find that the training world has little knowledge of formal assessment methodology, and most people outside of the testing industry don’t know what a psychometrician is.
People ask, “You’re a what? A psycho what?”
BM: You’re a psychologist.
DT: Well, that’s it, right? No, I’m not a psychologist. Although a lot of psychometricians are psychologists! They have PhDs in quantitative or cognitive psychology, and they’re developing what I refer to as more the “touchy feely” kinds of measurement instruments, measuring abstract constructs such as empathy, joy, personality traits, etc. There’s a lot of interesting work in this area and the measurement science is the same no matter what you are measuring.
Skills and competency testing is more concrete and it’s easier to teach people how to do this well. We’ve focused on moving away from “knowledge checks,” and improved our ability to measure application of skills.
BM: Yes, love that.
DT: We’re not testing knowledge, right? The 5 Moments of Need (5 MoN) approach is about performance support. We’re building competency aligned to performance objectives. We want them to DO vs. KNOW.
BM: Yes!
DT: I feel like the 5 MoN EnABLE methodology aligns so well with the principles of competency testing.
BM: It’s amazing and wonderful and was leading the horse to water for me. I was an L&D professional first in certification. I was part of the certification group at Microsoft, and it was interesting. The tail would wag the dog sometimes. There were those who bought our product or were into other areas of learning, and those that came and asked for training—but they weren’t really asking for training. They were asking for performance, right? We taught them to ask for training but that’s the means to an end. So, when I found the 5 MoN EnABLE methodology with Dr. Con Gottfredson, it was like, “Where has this been?!” Because it aligns so much to what I think we’re called to do and therefore can measure. It’s just such a shift for many folks who are steeped and have a high pedigree in L&D.
Let’s get into the idea of “levels” of measurement. Kirkpatrick came along and I was fortunate to meet him a number of times. I really appreciated him and his son, who continued his father’s work. I’m also very close to Jack and Patti Phillips, and the efforts they’re trying to make in their work, and I think there’s a misunderstanding or lack of appreciation for the shades of gray in measurement. I remember talking to one of the folks I just referenced about these shades of gray and they said, “The danger of numbering something is that you imply you know better than the other, or if one can go higher in the numbers then the other ones are useless or you shouldn’t do them, etc.”
So today, I’d love to bring clarity to Level 1 (L1) data.
DT: Yes, let’s.
BM: What do we mean by L1 data?
DT: I’ll be honest. I think Kirkpatrick was valuable for many years in providing some guidance to organizations that may not have otherwise known what to do about evaluation. But I think when workforce training is focused on performance support (as opposed to formal training), Kirkpatrick doesn’t align. This break allows us to reconsider the value of L1 data, or learner sentiment data. If the experience isn’t formal learning, what data about a learner is valuable?
Not to belittle learner satisfaction, but as we move further away from formal training, we’re being strategic about when and where we ask satisfaction or sentiment questions. It is still very important in formal learning settings to ask questions about learners’ experiences.
I think it’s always important to stop and say, “Why are we asking these questions? Are we doing it because everything must have an L1, L2, L3, and L4 measure? Or are we doing it because we think the sentiment of the people who are consuming the solution we’re providing is really important to telling our narrative and that it will help improve learning and job performance?”
BM: Sure.
DT: When someone says, “Diane, I want to do a survey analysis,” I ask, “Why? What’s the purpose? What’s the desired outcome?” Most of the time, they want actionable feedback. Well, surveys aren’t the best way to get actionable feedback. If you want information that improves learning outcomes, sometimes it is better to use focus groups or conduct interviews.
If you want quick insights into what’s working and what isn’t, or to what extent something is working, then surveys are a reasonable approach. Sometimes we do this because a number is more compelling to leadership than the insights from feedback.
Good survey design starts with understanding what the purpose is and what claims you want to make based on the data. In addition, you need to capture clean data from a sufficiently large, representative sample that supports the learning and performance narrative. However, I think we always have to keep in mind that it’s sentiment—people’s feelings—and what I feel in this moment might change five minutes from now. I may like it today but tomorrow I don’t. If you’re conducting a week-long training, sentiment may vary from day to day and we’re usually only measuring at one point. So, what is it you really want to know and is this the best way and time to capture it? For example, surveys are a reasonable way to measure whether a learner feels confident or enabled to do their job based on an experience, but we can’t say that they ARE enabled to do their jobs based on those results.
BM: I love that, and I’d love your opinion on something, Diane. There are always the infamous questions, “Do you feel like you can do your job? Can you apply what you’ve learned here on the job?” What’s your feeling on those questions (or similar questions) relative to what we can or should gain out of those? And maybe the answer is “not a lot,” but those are found so often in these surveys.
DT: Right. I will usually guide people to ask questions like, “Do you feel the content was relevant to what you’re doing?” Asking if it’s “relevant” is preferred because otherwise you’re asking people to make predictions before they use what they’ve learned. If you really want to know whether it helped them do their jobs, determine the right timing and/or cadence for the survey so that you’re measuring after they’ve had a chance to apply what they’ve learned. Of course, that just applies when we’re doing more formal training (e.g., eLearning, ILT). We do a lot of formal training for onboarding, and I think probably a lot of organizations still do when it’s for highly complex, highly critical work.
It may be easier to ask those kinds of questions about whether people feel they can do their job based on a performance support solution. Let’s think about what you put into a Digital Coach. If you use polling in a Digital Coach or other tool, the polls are usually so short that psychometrically we don’t consider them a reliable measure of satisfaction. However, polling can give us a lot of really good information about in-the-moment experience as someone is accessing and using the support tool—to say, “That was really useful.”
We often find this sort of polling on web pages. In fact, I recently asked Meta’s AI what the research says about polling on websites. Like all sentiment survey results, they’re mixed. I haven’t dug deeply to see what makes something (in this case a website) positive (i.e., in what situations did you get something helpful out of those measures vs. the situations where you don’t). I think short polling provides no more than a signal. In general, you get a lot of positive responses unless there’s something wrong, or people won’t respond at all unless something goes wrong on a web page or within a tool. You end up with two extremes.
BM: I want to keep on these themes of confidence. You hear a lot of talk about confidence measures. In your opinion, how does that play into this kind of data analysis or data collection?
DT: I think these are important questions to ask, but people writing surveys rarely want to make them long enough to get truly valid, reliable results. If what you really want to know is level of confidence, then focus all your questions on confidence. Generally speaking, we don’t want fewer than three to five questions aligned with a theme.
A lot of surveys ask very distinct questions, so you’re not getting a strong signal for any one theme. If you can narrow down the general themes such as confidence, engagement, usefulness of materials, etc., you’ll get a stronger signal. If I really want to know if someone is confident, I’ll ask questions about whether they are confident in their ability to use the tools we’ve given them, if they are confident in their ability to apply new knowledge and skills, etc.
I always recommend starting with a framework, so you have a clear plan for what you are measuring. What do I expect people to get out of the training? What results would support the expected outcomes and the learning narrative? Then I categorize and make sure I’m asking at least three questions per category. If you cannot, for practical reasons, ask three questions, I encourage those using the results to keep in mind this limitation when interpreting the results.
BM: Let’s talk a little bit about methodology here first, because you’ve mentioned polls. What are the tools of this trade? I think we myopically focus on surveys. We think if we’re going to do an assessment, we’re going to do an end-of-class survey and that’s kind of the hammer for the nail. Can you broaden our view a bit about ways in which we can orchestrate and deliver these things?
DT: The methodologies that I have actually used for L1 data collection are all some variation on a survey, with some qualitative collection with focus groups or interviews. If you can pull together a very representative focus group, you’re going to get far more valuable information for program improvements, but it doesn’t satisfy leadership’s desires for numbers, nor does it allow you to track trends over time. Alternatives to surveys are polls, which are just shorter surveys delivered “in the moment”, including something as simple as thumbs up/thumbs down with a field for text feedback. I’ve suggested this in a few cases where we just want to know if something went wrong.
BM: So, what’s your feeling about timing relative to polls/surveys? You mentioned answers reflecting sentiment during the moment in time when the questions are asked. We often hear the debate about 30-60-90 days. What’s your feeling about cadence? What’s your feeling about the tool that should be used within that 30-60-90 cadence? Same tool? Same questions? How do I use time to my advantage?
DT: It depends on your scenario. If I want to know whether people’s sentiments have changed, I’m going to do more of a pre/post approach using the same questions. If I want to see whether their sentiments are changing over time—remember we talked about wanting to know if training was helpful for someone to do their job (e.g., did the training enable you to do your job better?)—then I might deliver the survey more than once. In this case, we would survey them the moment they finish the training and then survey them again later. I would say 30-60-90 gives you a rule of thumb, but I think that really depends on what the job role is and how much they’re using what they learned.
I haven’t seen the research on surveying people again after another 60 or 90 days. I would question whether it’s worth continuing to collect data unless it’s a very specific scenario where you’re getting them to an entry-level point and they have to build up expertise, so you’re expecting changes over these periods of time. In that case I’d say, “What is that period of time in which you’re expecting a change and are there continued interventions?” For example, there may be performance support that is enhancing performance, thus you’re no longer measuring the impact of the initial intervention. The further out you go, there are more and other variables that impact their sentiment about what they learned.
BM: I love the variability of the role, right? I mean, if I’m required to do professional development or performance reviews, to your point, if I took training in January and I’m on a calendar cadence that’s going to hit me at weird times of performing vs. understanding the ebb and flow of the workflow, how do we time training closer to performance or the potential of getting better at something based on how frequently I do it?
DT: The response rate is also affected by passage of time. It gets to that difference between delivering an assessment after a solution’s been implemented, like a formal training event, vs. asking everybody to complete an assessment at certain points in the year because they’re doing ongoing activities. You’ll get more targeted feedback when you are strategic about timing.
BM: Let me ask a couple quick things. I think some of these are just fundamental, but at least in my journey, the answers have been all over the map.
First, I’d love to hear from somebody with your experience and pedigree about scale size. Do you have any principles or guidelines for that? This is a notorious question.
DT: I think you have to take a number of things into consideration. Think about who your audience is and the cognitive load. If you’re going to have 30 people respond to a survey, don’t do 7-to-10-point scales because you’re going to end up with too much missing data (no one is selecting certain points on the scale), resulting in a problem with your scale maintaining its ordered structure. Now we’re getting into fun psychometrics stuff.
BM: I love that!
DT: In my experience, 5-point scales are most common. With anything below that, we’re not getting enough variability to produce reliable results; although I would say 5- or 6-point scales because you don’t always want a neutral point.
BM: Right. I love that because I think we assume we should have 3 out of 5.
DT: I think it depends on what you’re measuring. For example, we use surveys to help us determine how critical or how complex something is for Rapid Workflow Analysis. That is then used to create assessment blueprints (a framework that specifies what to assess). In those cases, if you are a subject matter expert and you do this for a living, I don’t want you to be neutral. You need to have an opinion about this. In that situation, I don’t recommend an odd number of points.
BM: It pushes them either way, right? They can’t sit on the fence.
Next question: AI. It’s in everything, and the devil is in the details in terms of changing our industry. If I’m an L&D professional not knowing what to think, what are your perspectives?
DT: For me personally, AI is changing my role because I work for an AI company. We’re embracing that and I love it because at this point in my career, it would be very easy for me to become complacent, so I love that I’m drawn into this.
Industry-wide, I don’t know as much about AI’s impact on L1 data, but I know that in assessments and psychometrics we’re quickly exploring, testing, and implementing—but with caution. A lot of what we do in the assessment industry involves concerns of legal defensibility and fairness. There are a lot of privacy considerations. But for companies that are trying to build out some measure of sentiment, I always start with AI. It’s my personal research assistant. I love it. It makes my life so much easier. When I get a request to design a new measurement instrument, I outline the construct to be measured with what I know, then prompt AI to report what’s in the research. Once I have a clear design and framework, I prompt AI to provide a specified number of sample questions or statements that align with my framework. For example, “Here are the three things I want to measure. This is what I want to get out of these measurements.” In this case, I may ask AI to give me 10 examples of how to determine the level of confidence individuals feel based on the training or performance solution we’ve provided. Then I review and choose the ones I like and make edits. I may also provide additional guidance to clarify what I want to measure. The result is usually a combination of AI’s suggestions and my own additions.
Keep in mind that AI hallucinates, so we always want to be really careful. It definitely hallucinates articles, so I always double check citations. It can also draw inaccurate inferences from data, so I use it more to help me determine how to analyze something or write formulas and code. I also use it for qualitative data analysis, identifying initial themes and patterns, such as text responses to surveys.
AI is very useful with text analysis.
BM: I love the words you said: “guidance” and “prompting”. I’m finding that people struggle with the idea of conversation vs. a question when it comes to AI. They’ll ask for it to create a survey on something, indicate they want it to have five questions, and then when the five questions are produced, they think, “Well, OK. It got me started. This is great.” And then they leave it there. I don’t think we’ve learned yet the intellectual (if that’s the right word) side of the prompting and guiding and guardrails. We don’t know that we can say, “I like questions 3 and 9. Can you reword 2, 4, and 6?” That’s our responsibility, but I think the heavy lifting we continue to shoulder vs. conversing with AI to get closer to what we need is something we’ve got to improve upon, and I would love to see what you’ve done in that sense because of your understanding of the science. I’m sure it’s just remarkable.
DT: The important thing is to understand the methodology in the first place and how to design a measurement instrument. Understand what you want to measure. Make sure you’ve explained that to the AI. I never ask it to just write a survey. I always say, “Give me sample questions based on what I want to measure in these three categories.” Then I build that instrument correctly. Of course, I’m a psychometrician and I’m a control freak, so I’m not going to just let AI do it, but there are no guidelines that say you can’t ask more questions, right?
BM: It’s remarkable how more guiding and trusting are required, but I see people stop and either get cautious or think that’s as far as they dare go.
My final question for you is about pitfalls. If I’m dipping my toe in L1 data, can you name a few things you see most people do incorrectly?
DT: Number one on my list is using L1 surveys to evaluate instructors. It’s generally not a fair, reliable, or valid measure of instructor effectiveness, especially if you’re going to ask one question about the instructor within a learner satisfaction survey. If you are going to ask about instructor effectiveness, define what good instruction looks like and ask multiple questions. Basically, my advice is to put some psychometric rigor into it—because people’s lives can be impacted by these measures.
Number two is understanding the scales. For example, understanding that if the ordinal structure of a Likert-type scale doesn’t hold, then your results aren’t meaningful. A rating of five should represent higher sentiment than a four, and so on and so forth. However, people can interpret the points differently, particularly if you only anchor the end points. This is less common with shorter scales and more of an issue with longer scales, such as a 10- point scale where we only label the end points. To avoid issues, avoid longer scales unless you expect larger sample sizes and use clear labels and/or descriptions aligned with score points to reduce subjectivity.
You also need sufficient sample sizes. People always ask, “What does that mean? How many people do you need?” There’s not a single answer to that as it depends on the size and diversity of the target audience. More important than numbers is representation. Do the people responding represent all the facets of your target audience?
BM: Wow. Such good stuff. I’ve had five or six “Aha’s” in this conversation with you today. Thank you so much!
DT: Well, I’m equally grateful to have the opportunity to work with you and Con. I can’t tell you how much it has expanded my horizons as a psychometrician.
BM: That’s a great compliment.
DT: It really has been incredible. I’m a great proponent of 5 Moments of Need. I think it’s an incredible model.