Ed. Magazine Pilot, Test, and Assess Ideas First Professor Tom Kane explains why schools should measure what works before scaling new programs Posted September 25, 2026 By Lory Hough Education Reform Evidence-Based Intervention Student Achievement and Outcomes Teachers and Teaching Professor Tom Kane Photo: Veasey Conway/Harvard University Pilot, test, assess. It's not the obvious answer to the question: How do we help students succeed? But as far as Professor Tom Kane is concerned, it’s one of the best answers. When asked why, he says, “Because student behavior and teacher behavior are really complicated and it’s impossible to anticipate every unintended side effect. It’s impossible to anticipate how people are going to respond when you impose a cell phone ban or you propose a change in the teacher evaluation system. We have ideas about how people will respond, and that’s what leads people to advocate for these kinds of policies, but actually we don’t know. I have multiple times seen things that I was really optimistic would work that turned out not to work.”One example happened a few years ago when Kane and his colleagues at the Center for Education Policy Research at Harvard University (CEPR), which he founded, did a study of elementary math textbooks. Prior studies suggested that certain textbooks would be more effective than others. CEPR collected data from textbooks being used in six states and measured differences in student math skills. Kane assumed they’d find big differences and then redo the study annually to assess if that held. “We worked on that for weeks and weeks and weeks before we decided there are few, if any, differences in student achievement gains for people using different textbooks,” he says. As a result, they didn’t repeat the study.Another project they took on in 2005, the year CEPR started, was in collaboration with Joel Klein, then-chancellor of New York City Schools. “Joel had made a big bet on the New York City Teaching Fellows program. They were recruiting a quarter of their new teachers through this alternative route program,” Kane says. Klein asked CEPR to evaluate the program, which was created to train professionals and recent college graduates to teach in high-need, low-income public schools. Klein was optimistic that CEPR would find big differences between the teaching fellows and other novice traditionally certified teachers.“And that’s not what we found,” Kane says. “Actually, there was little, if any, difference once we controlled for students’ baseline achievement. There was little, if any, difference in the efficacy of the average teaching fellow versus the average traditionally certified teacher. That was not something Joel Klein or we were expecting.” To his credit, Kane says, Klein “basically shifted on a dime” and said, “Oh, OK. So this isn’t about recruitment. It’s about assessing people’s performance on the job in the first couple of years of teaching and making sure we keep the really great teachers.” It’s these kinds of unexpected results that make piloting, testing, and assessing critical in education for student success, Kane says. “There are ideas that sound good on paper, but it’s just impossible to anticipate all the ways in which it can go wrong,” he says. “Think about your home improvement projects. The way you think something is going to work never does, but what’s different is that in a home improvement project, you get to see that it’s not working because the lamp won’t turn on or the paint job looks horrible. In the case of education, we don’t get to see that our ideas don’t work because we don’t always have comparison groups. We see, oh, test scores went up or test scores went down, but that may be happening for reasons that have nothing to do with what we just did. And the only way we know that is if we have a comparison group, if we tried it in some classrooms and not in other classrooms and then compared.” This doesn’t happen in education as much as it should, he says, because practitioners can be “overconfident” that their great ideas for improvement will work. “It’s just human nature to be overconfident. That’s sort of built in,” he says. And unless you have a conversion experience, like Kane has had over the years with projects like textbooks and teaching fellows, “it’s very hard for people to let go of the assumption that this idea you have will work.” "It's impossible to anticipate how people are going to respond when you change the curriculum or when you impose a cell phone ban or you propose a change in the teacher evaluation system." Professor Tom Kane Luckily, says Kane, it’s easier these days to pilot, test, and assess ideas and assumptions because more data is available to use. In the past, he says educators needed a Ph.D. or an expertise in something like program evaluation to figure out what was going on with students. But since 2005, robust state data systems have become widely available, making it easier to compare data points, such as the achievement gains of students who get a given intervention against those who don’t. “They’re tracking students over time,” Kane says. “Actually, it’s part of the reason why we started CEPR. We saw that there was now a chance to learn a lot more.” This emphasis on student data collection came out of the No Child Left Behind Act, which said states had to start testing in grades three to eight. “After a few years, this made it possible to follow student progress.” Kane stresses, though, that just having more data isn’t enough — it has to loop back to pilot, test, and assess and these actions need to become habits. “The problem is [some educators] haven’t taken the extra little step to say, OK, we have this data and we know which kids are, for example, getting tutors,” he says. “Why don’t we put these two things together and see if the kids who are getting tutors are seeing faster gains than the kids who are not getting tutors? Let’s not just assume that tutoring must work.” In fact, he says, using the tutoring example, districts spent a lot of money after the pandemic on tutoring, assuming that high-dosage tutoring was the answer to making up lost learning. The assumptions were based on pre-pandemic research that suggested that high-dosage tutoring would lead to large gains. “Lots of people were arguing, OK, this is how we’re going to rebound from the pandemic. We’re going to find the kids who are furthest behind and assign them tutors,” he says.CEPR started working with several school districts that poured money into tutoring, only to discover that gains were small. “It turns out there were all sorts of implementation problems. Rather than three sessions a week for a whole school year, or 108 sessions, which is what the pre-pandemic research was based on, we were discovering kids were usually often attending for five or 10 sessions total.” Kane said the idea behind high-dosage tutoring wasn’t necessarily bad, but the implementation of it often was. “Many districts were walking around thinking, ‘We’re solving this problem. This worked. All this prior research told us it was going to work. We’ve done our part.’ But they didn’t bother to track or ask, how many sessions are kids getting? What are the achievement gains the kids are getting? Instead, it was, ‘Well, we spent our money on the tutoring, so now let’s work on the next problem’ without realizing what you tried didn’t work. The thing that was missing was they didn’t track the data to compare what happened to the people who got the tutors versus the people who didn’t. If they had, they would’ve learned, as the districts that worked with us learned, ‘Gosh, many fewer kids than we thought are showing up for these tutoring sessions.” Using the retail sector as an example of a sector that has figured out the importance of tracking data, Kane explains that checkout scanner software eventually made it easier for stores to gather data on every purchase — how many units of fuji apples sold versus macintosh — and they no longer had to “guess” at what worked. “All of a sudden people said, ‘We don’t have to make these big bets anymore because we’re seeing that some of our ideas are not working,” he says. “We’re doing these market tests and something we’re pretty sure was going to work, turns out not to work. Now everything needs to be tested in retail before it’s scaled up. Walmart doesn’t scale up anything across the chain anymore without doing a pilot test. “That hasn’t happened yet in education, but we’re close,” Kane says. “We’ve got the equivalent of the transactions database in student test scores. What we lack is just that last little bit, which is using those things to do pilot tests. And honestly, that’s what we’re trying to do at CEPR. We’re trying to provide that last little bit.” Ed. Magazine The magazine of the Harvard Graduate School of Education Explore All Articles Related Articles News New Education Scorecard Finds “U-Shaped Recovery” High- and low-income districts improve most since 2022, while middle-income districts (30–70% federally subsidized lunches) lag News U.S. Learning Rebound Underway According to Education Scorecard Professor Tom Kane discusses new data showing gains in math and reading, even as challenges like absenteeism persist News Short-Term Education Recovery Effort Shifts to Long-Term Reform As chronic absenteeism slows the pace of academic recovery, researchers urge states and districts to recommit to effective interventions