Tuesday, April 1, 2025

Betts, Julian R. & Loveless, Tom. (Eds.). (2005). Getting choice right: Ensuring equity and efficiency in education policy. Reviewed by Yun Teng, Arizona State University

Education Review. Book reviews in education. School Reform. Accountability. Assessment. Educational Policy.

Betts, Julian R. & Loveless, Tom. (Eds.). (2005). Getting choice right: Ensuring equity and efficiency in education policy. Washington, D.C.: Brookings Institution Press.

255 pp.
$49.95 (hardcover)   ISBN 0-8157-5332-2
$19.95 (papercover)   ISBN 0-8157-5331-4

Reviewed by Yun Teng
Arizona State University

December 22, 2006

Although much of the debate over school choice has been focused on whether school choice is “good” or “bad,” Getting choice right: Ensuring equity and efficiency in education policy endeavors to move beyond polarized debate and draw attention to practical concerns about the policy design of choice proposals. The book is a product of the 2001 National Working Commission on Choice in K-12 education initiated by the Brookings Institution. It is a collection of papers from different scholars studying school choice. The contributors view school choice as an accepted and growing part of the nation’s education policy concerns and focus on how the benefits can be maximized and risks mitigated. The first part of the book discusses from a policy perspective how to promote a vigorous demand for and an adequate supply of education choices. The second part examines the impact of choice on the distribution of education opportunities.

In Chapter 2, Betts explores how economic theory can inform choice design. He discusses the benefits of consumers from markets that are on a continuum: from “perfectly competitive” to monopolistic, with perfect competition providing the most efficient market. He then discusses the assumptions required for perfect competition and how policies can address violations of these assumptions. The basic argument is that what determines the outcome of school choice are the relative amounts of resources available to parents initially. Betts makes two proposals: quotas combined with lotteries for school admission, and an inter-school tradable market of high-achiever enrollment rights to address possible unequal educational quality resulting from competition.

Hamilton and Guin analyze the demand for school choice by examining how families make choices in Chapter 3. They draw the conclusion from previously published research that academic quality is parents’ foremost consideration, but caution that parents have various choices that can only be understood in their diversified contexts. Parents might stress different sets of values, leading to greater school segregation. Moreover, what parents say sometimes differs from what they choose. It is unclear whether parents report socially acceptable values to hide their true reasons, or whether they have insufficient information to make informed choices. Also unknown are the processes, channels and factors involved in parental choice, as well as the mechanism of the two-way information flow between families and schools. The authors call for widely disseminated information for parents and the incentives for parents to use the information.

In Chapter 4, Betts, Goldhaber and Rosenstock study the supply side of school choice. Having outlined a theoretical model, they analyze influences affecting the supply of both choice schools and teachers. They argue that major constraints on the supply of schools are market price and cost structure, which are significantly influenced by the political and regulatory environment. Their recommendations include increasing funding, deregulation, and reducing policy uncertainty.

Hess and Loveless examine the “black box” mechanisms producing the purported achievement effect for active choosers in Chapter 5. They argue that such studies are essential for the replicability of such an effect on a larger scale, especially given that school choice encompasses an array of different arrangements and that the effects of choice may be confounded with success or failure of particular practices at choice schools. They look into structure, population, and classroom instruction of choice schools; and discuss the possibility of diminishing benefits and limitations of school choice when implemented on a larger scale. They urge for caution against the wider implementation of choice and more research on possible difference in effects between choice plans adopted on small versus large scales.

In Chapter 6, Goldhaber and his colleagues analyze effects of choice on nonchoosers. Among their critical analysis of the possible benefits and detriments, particularly worth noting are the educational inefficiencies, regulatory costs, peer effects, underprovision of public education, and symbolic competition as possible negative effects of choice. Having emphasized the influence of contextual influences on choice dynamics, they urge that student flows should be tied with school resources. To solve the incentive-capacity dilemma— that is, the difficulty for schools losing enrollment to make improvements with less funding— they suggest that underperforming schools lose less than the marginal cost of educating the departed students to retain more funding for improvement in the short run. They also insist that all forms of public education funding be tied directly to per pupil spending in public schools to make explicit that the stake the entire society has in public schooling.

In chapter 7, Gill provides a conceptual framework for evaluating school choice’s effects on integration. He first analyzes the factors influencing integration in schools within a hypothetical system absent of choice and then explores the possible added-on effects of school choice. He then offers suggestions for policymaking and research. His emphasis is on the quality of integration, consideration of constitutionally protected choices based on residence and private schools. Gill prefers school-to-school or even classroom-to-classroom comparison between choice and traditional public schools. He also prefers a dynamic model to a static one to account for the effect of school choice on integration.

The effect of charter schools on integration in Michigan are examined in Chapter 8. Ross finds that the racial composition of charter schools is not dramatically different from traditional public schools in their vicinity. However, charter schools are significantly more likely to locate in districts with greater segregation in traditional public schools. Although this indicates that charter schools are responding to minorities’ educational needs, it exacerbates public school segregation. As the main effects occur when charter schools serve large proportions of public school population, Ross suggests putting a cap on the percentage of a district that charter schools can serve.

Henig surveys the political conflict over school choice in Chapter 9. He argues that although the four dimensions of education provision (delivery, financing, regulation and decisionmaking) can be analytically autonomous, they might not be empirically independent. He distinguishes pragmatic and systematic privatization. Given the fact that the U.S. education system already comprises both public and private characteristics and the deeply-imbedded ethos of incrementalism and pragmatism, concern about pragmatic privatization is overblown. However, caution is needed for systematic privatization involving shifts of power, public perception and institutional arrangements, all of which might weaken democratic control. Unlike most authors in this volume, Henig is less enamored with policy design, which is ultimately dependent upon the nature of political constituencies and the institutional capacity of government. That granted, it is likely that political power will trump pragmatism and neither proponents nor opponents will cede territory in pursuit of common ground.

In the last chapter, Wolf concludes from a review of existing quantitative research on choice schools in teaching civic values that private schooling and school choice rarely harm and often enhance such learning. He recognizes, however, that these studies are not without limitations. Unlike Hess and Loveless, Wolf argues that although precisely how schools of choice teach civic values remains hidden in the proverbial black box, the importance of the existing studies should not be discounted.

This single volume covers most issues in the current debates over choice. The shift from polarized battles to pragmatic concerns is to be appreciated. Contributors come from both camps. Although it is not difficult to identify their affiliations, the line in between is less clearly limned. The book itself signals a move toward “pragmatic privatization,” as in Henig’s chapter. However little the move is, it initiates a healthier discourse for future discussions.

There are certainly limitations. I share Henig’s caution that it is unwise to become obsessed by policy design. The belief that good policy design will “get choice right” is not much different from the belief that school choice is “the magic answer.”

We need to consider the various contexts of school choice. Although all contributors agree that the black box requires exploration, little effort has been made in this direction. Methodologically, unraveling these contingencies would be accomplished neither by quantitative (especially randomized) studies taken as the “gold standard” by some contributors, nor by constructing grand theories. Qualitative studies represent a more promising approach to answering these questions, but few have been cited or recommended as models for future research.

Paradoxically, many of the policy suggestions contradict the alleged advantages of choice schools and the alleged failure of traditional public schools. For instance, although choice schools have been extolled for their economic efficiency and political autonomy, many contributors call for more funding and deregulation. This is not to decry choice, but to remind the public that despite their differences, choice schools (as long as they remain “public”) and traditional public schools are in the same boat. On the one hand, the “necessary evils” are necessary for both types of school for good reasons. On the other hand, the difficulty choice schools now face mirrors the morass traditional public schools have been trapped in. The suggestions for “alternative” schools unintentionally reveal the fact that traditional public schools are not the scapegoat but the victim of the “education failure.” If traditional public schools cannot escape their fate doomed by the unfriendly economic, social, and political context, neither can choice schools.

All detours lead to the same destination. Historically, detours have been designed to bypass the contextual problems haunting U.S. education, but they have not led us anywhere else. Choice has its own values, but if it is meant to be just another detour, it will soon become clear that we again reached the same destination.

About the Reviewer

Yun Teng is a PhD student in the Division of Educational Leadership and Policy Studies in the Mary Lou Fulton College of Education at Arizona State University. His major research interest is in choice-based school reform.

Copyright is retained by the first or sole author, who grants right of first publication to the Education Review.

Mullen, Carol A. (2007). Curriculum leadership development: A guide for aspiring school leaders. Reviewed by J. Craig Coleman, Stephen F. Austin State University

Education Review. Book reviews in education. School Reform. Accountability. Assessment. Educational Policy.

Education Review/Reseñas Educativas/Resenhas Educativas

Oppenheimer, Todd. (2003). The Flickering Mind. New York: Random House.

528 pages
$15.95   ISBN 0-8129-6843-3

Reviewed by Andrew J. Rotherham
Education Sector

May 15, 2006

For another review of this book
see the review by Groff (2005).

Todd Oppenheimer is not a Luddite. That it is important to mention a caveat like that in a discussion of educational technology speaks to the pervasiveness of technology and the strength of our collective faith in its transforming potential. It is an important caveat though because Oppenheimer has produced a 400-page indictment of one of the most popular (and expensive) school reforms of the past twenty years: The massive effort to put computers and technology in the nation's classrooms. Seeing him as a reflexive opponent of technology would make it too easy to dismiss his important observations out of hand.

In The Flickering Mind Oppenheimer walks readers through a history of computers in school, the bold promises about technology’s potential, some case studies about various communities, and some general ideas about ways he thinks schools must now reform the technology reform. Along the way he offers some penetrating insights and interesting reporting about the impact of technology on education and he sounds some useful cautions. However, though I share many of Oppenheimer’s concerns and share many of his conclusions, The Flickering Mind is in places a frustrating book because it reads with the confidence of an analysis that is easier in hindsight than at the messy inception of an idea. But this gets ahead of the story.

The idea of technology in education is nothing new. The current generation of proponents of e-learning, distance learning, online classes and so forth stand on the shoulders of previous generations who have tried to improve on the basic teacher to student relationship. Oppenheimer notes this and gives a brief summary of the various educational promises made on behalf of “technology” long before computers were anything more than the imaginings of futurists and dreamers. He recounts the enthusiasm of Thomas Edison for motion pictures as a way to revolutionize teaching and later similar predictions from others about radios. He also discusses the federal government’s first forays into technology, which would later grow into major programs and initiatives.

His reporting is vivid. Oppenheimer goes inside schools in Harlem, rural West Virginia, Napa, California and Maryland's affluent Montgomery County to paint a fascinating picture of their experiences with technology in classrooms. This is no mean feat. Writing about what happens inside schools with texture and perspective is a challenging task, especially when policy issues are also involved. Oppenheimer succeeds in conveying both an account of the experience of these schools and how it relates to his larger narrative.

He takes us to Hundred, West Virginia, the sort of small isolated community that does not jump to mind when thinking about the dot.com revolution. But Hundred High School has many of the whiz-bang gadgets that you would expect to find in any latte soaked Seattle company. Yet it is no panacea. Oppenheimer writes about a math class, graphing equations using computers. The computers enable the students to work faster and plot more graphs and tackle more problems in class as a result. These are ostensibly exactly the sort of enhancements that proponents of technology seek. But, before they get to work, Oppenheimer reports that it takes the teacher 20 minutes of valuable class time just to get all the students ready to use their computers for the lesson. Perhaps in the end trade-offs like this are worth it: the additional time spent getting things in order is outweighed by the increased productivity the technology will generate over time. Oppenheimer doubts it, no one knows for sure, but everyone on all sides of the technology debate can agree that there are trade-offs. It's just one example of how technology is not an absolute blessing in the classroom.

In New York City, Oppenheimer relays the frustration of teachers who have computers and other gizmos dropped in their classrooms and are unable to use them, get support or service and ultimately end up frustrated with what amounts to overpriced paperweights cluttering up their space. He argues that rather than blaming the teachers, or just calling for more training for them, these problems are symptomatic of deeper problems with the current approach to technology in schools.

In the end, these and other problems convince Oppenheimer that the current tradeoffs are not balance positive for students and schools. Aside from the distractions and shortfalls of much of the hardware, he argues that even when everything is working like it should the internet is too often merely a data dump of information some valuable, some incomplete, some biased or misleading, and some flat out wrong. The problem, as he sees it, is that like other media sources a critical view is necessary for active and informed consumption. He worries that the passive nature of the internet experience for many students, and the reliance on technology rather than teaching that often comes with it, is dulling these skills if imparting them at all.

Those promoting technology, whether in government or industry, come in for particular criticism for fostering this state of affairs. Oppenheimer chastises Clinton Administration officials for their unbridled enthusiasm about technology and many of their efforts to expand its reach in schools. To be sure, various industries stand to gain or lose a lot depending on government policies and companies promoting classroom technology are no exception. Consequently they're not shy about pushing their agenda with those in government. But technology’s biggest boosters in the administration and the public sector in general are not craven, opportunistic, or deluded by lobbyists. Rather they are motivated by the best of intentions.

And that’s a big part of the problem, Edison was sincere too. Like many educational policy issues, this one boils down in no small part to a clash between good intentions, practical realities, and what historians Larry Cuban and David Tyack have called “the grammar of schooling,” namely the enduring characteristics of education that are slow to change.

Some of the educational problems that Flickering chronicles are neither unique to nor caused by technology, and in places the book does read like a fault finding mission made easy through the lens of hindsight. Still, there are trenchant observations in Flickering with important implications. Oppenheimer notes that it is entirely possible that some of the inequities that plague our schools will only be exacerbated not lessened by technology. This is not the well known concern about a “digital divide” in access, but rather a subtler problem of differences in the curriculum, teaching, and learning between affluent children and other children. After all, despite the prevalence of technology in the lives and work of upper-middle class professionals, the kind of schools they most often seek out for their own children tend to emphasize small classes, personal connectedness, and quality interactions between teachers and students.

But technology has reach in education beyond the classroom. Unfortunately, Oppenheimer does not deeply examine its potential outside the classroom. Here is where technology may help us address larger issues. For instance, the venerable financial analysis firm Standard and Poor's has developed a data platform that allows policymakers, and the public, to examine educational performance, expenditures, demographics, and other variables. Just a few years ago, this sort of analytic leverage was a pipe dream for wonks and analysts. Likewise, on-line assessments are showing early promise to lessen the amount of time devoted to student testing, get results back to teachers faster, and offer greater sensitivity than today’s pencil and paper tests. Over time, initiatives like these, which complement teaching rather than seeking to fundamentally change or replace it, may well have a more lasting impact on what happens inside classrooms.

Flickering should not be read as a determinative verdict on the potential of technology in schools but Oppenheimer’s cautions are worth heeding as the educational technology juggernaut goes forward. Besides the issues he raises, a tempered enthusiasm and critical eye are healthy considering that, for all its promise, technology alone cannot solve some basic problems in education, for example the “preparation gap” that starts poor children off in school on an uneven footing or funding disparities between different states and school. In terms of classroom applicability it's regrettable, though, that Oppenheimer largely stays away from making recommendations of his own beyond generalities and related educational issues such as funding and teacher pay and respect. Those issues are not unimportant, and he's right in noting that educators are bombarded with plenty of recommendations and advice as it is. Yet on the heels of a critique like this it seems like a cop-out.

In the end, Oppenheimer is right that technology in schools hasn't been an unbridled dot.com boon. Nor, however, has it been a dot.com bust and the potential remains impressive. Thankfully, not unlike shrewd investors in the dot.coms, savvy educators are becoming more particular about technology and are internalizing the lessons of the recent past. Like most things, progress here will be messy and full of mistakes and false starts, but despite the problems there are signs of progress nonetheless. Today we shake our heads when we consider that the ubiquitous Blackberries have more computing power than the spacecraft that took Americans to the moon. A generation hence we may look back on the infancy of classroom technology with similar bewilderment.

About the Reviewer

Andrew J. Rotherham is co-director of Education Sector (www.educationsector.org), a senior fellow at the Progressive Policy Institute, and a member of the Virginia Board of Education. He writes the blog Eduwonk.com. Rotherham previously served at The White House as Special Assistant to the President for Domestic Policy. He managed education policy activities at the White House and advised President Clinton on a wide range of education issues including the reauthorization of the Elementary and Secondary Education Act, charter schools and public school choice, improving educational options for disadvantaged students, and increasing accountability in federal policy. Rotherham also led the White House Domestic Policy Council education team, the youngest person to have done so.

Copyright is retained by the first or sole author, who grants right of first publication to the Education Review.

Alsup, Janet. (2006). Teacher Identity Discourses: Negotiating Personal and Professional Spaces. Reviewed by Elizabeth Smolcic, Pennsylvania State University

Education Review. Book reviews in education. School Reform. Accountability. Assessment. Educational Policy.

Education Review/Reseñas Educativas/Resenhas Educativas

Bracey, Gerald. W. (2006). Reading Educational Research: How to Avoid Getting Statistically Snookered. Portsmouth, NH: Heinemann.

188 + 20 pp.
ISBN 0-325-00858-2

Reviewed by Darrell L. Sabers
University of Arizona

May 29, 2006

Reading Educational Research: How to Avoid Getting Statistically Snookered is the latest book by Gerald Bracey. He delivers what is expected by those of us who read his work, with many examples of how others lie with statistics and interpret data to fit their agendas. The book is focused on 32 “Principles of Data Interpretation” that should provide many readers with a set of instructions to examine and understand reports and claims about various educational issues. The 32 principles are as follows:

Principles of Data Interpretation

  1. Do the arithmetic.
  2. Show me the data!
  3. Look for and beware of selectivity in the data.
  4. When comparing groups, make sure the groups are comparable.
  5. Be sure the rhetoric and the numbers match.
  6. Beware of convenient claims that, whatever the calamity, public schools are to blame.
  7. Beware of simple explanations for complex phenomena.
  8. Make certain you know what statistic is being used when someone is talking about the “average.”
  9. Be aware of whether you are dealing with rates or numbers. Similarly, be aware of whether you are dealing with rates or scores.
  10. When comparing either rates or scores over time, make sure the groups remain comparable as the years go by.
  11. Be aware of whether you are dealing with ranks or scores.
  12. Watch out for Simpson’s paradox.
  13. Do not confuse statistical significance and practical significance.
  14. Make no causal inferences from correlation coefficients.
  15. Any two variables can be correlated. The resultant correlation coefficient might or might not be meaningful.
  16. Learn to “see through” graphs to determine what information they actually contain.
  17. Make certain that any test aligned with a standard comprehensively tests the material called for by the standard.
  18. On a norm-referenced test, nationally, 50 percent of students are below average, by definition.
  19. A norm-referenced standardized achievement test must test only material that all children have had an opportunity to learn.
  20. Standardized norm-referenced tests will ignore and obscure anything that is unique about a school.
  21. Scores from standardized tests are meaningful only to the extent that we know that all children have had a chance to learn the material which the test tests.
  22. Any attempt to set a passing score or a cut score on a test will be arbitrary. Ensure that it is arbitrary in the sense of arbitration, not in the sense of being capricious.
  23. If a situation really is as alleged, ask, “So what?”
  24. Achievement and ability tests differ mostly in what we know about how students learned the tested skills.
  25. Rising test scores do not necessarily mean rising achievement.
  26. The law of WYTIWYG applies: What you test is what you get.
  27. Any tests offered by a publisher should present adequate evidence of both reliability and validity.
  28. Make certain that descriptions of data do not include improper statements about the type of scale being used, for example, “The gain in math is twice as large as the gain in reading.”
  29. Do not use a test for a purpose other than the one it was designed for without taking care to ensure it is appropriate for the other purpose.
  30. Do not make important decisions about individuals or groups on the basis of a single test.
  31. In analyzing test results, make certain that no students were improperly excluded from the testing.
  32. In evaluating a testing program, look for negative or positive outcomes that are not part of the program. For example, are subjects not tested being neglected? Are scores on other tests showing gains or losses?

After 20 pages of introduction of topics about data-driven decisions and abuses of data, the principles are incorporated into the text with examples and explanations. Many of these examples are based on reports of educational evaluations, and Bracey does not hesitate to include controversial topics with political agendas. Most of this book is very good reading, and it is not intended only for the reader who has no background in reading research reports.

The principles of data interpretation are, for the most part, very good points to remember when reading reports. I take issue with a few in terms of the wording, and have a problem with the organization. These problems are not important enough to relegate the book to the “wait for the revision” category, but might help with some readers focus better on related issues in the principles. For example, #s 4 and 10 relate to comparing groups, and #s 9 and 11 relate to types of scores being compared. One looking at these principles might have a difficult time understanding the order. Principle #18 is incorrect as described below, and many could have been improved by more careful wording (#s 18, 19, 28). For example, with the advanced placement tests we don’t expect all children to have been exposed to all the desired curricula (#19, which orders words differently from #20), and in #28 the word “interpretations” would be better than “statements.” But this need for minor changes should not change the overall rating of excellent for the set as a whole and as a basis for the book.

The treatment of topics covered in the principles is mostly excellent, and the reading of the book should enhance the preparation of a reader to become a credible consumer of educational research. If the readers of most of what is published about testing used these principles, the writers would be encouraged to prepare their reports more carefully.

There are many reasons to support the claim that this is a very good book. However, for this review, the invitation that Bracey presents on page 172 as he closes his book is an opportunity that must be accepted. He states, “I would be happy to hear from any reader suggestions for what a second edition of this book might add, delete, alter, or reorganize.” That is a challenge any reviewer should be eager to answer.

OK, Jerry, here are some comments that your fan Sabers (whoever he is) would like you to consider in your second edition (and hopefully the reader of this review will benefit from these suggestions as well).

When discussing principle #3 Bracey discusses salaries of teachers “on average”; however, principle #8 tells the reader to “make certain you know what statistic is being used when someone is talking about the ‘average’.” Now if a critical researcher like Bracey uses “average” without telling the reader what statistic is meant, why would the reader be expected to practice the necessary diligence to “make certain…?” This particular sloppy use of the term “average” also detracts from a thoughtful explanation of principle #3, “look for and beware of selectivity in the data.”

A list of variables is presented on pages 38-39, but two of those are “student surveys (of former as well as current students)” and “changes over time in all of the above variables.” Now neither of these is a variable, and that is not a good example for an author to use when teaching the reader to be a careful consumer of research.

One would hope that the error in reporting variables would not distract the reader from the following sections on dealing with rates, numbers, and scores, and dealing with comparison of rates or numbers over time (make sure the groups remain comparable as the years go by). This is an excellent part of the book, but I think it would have been improved if the introduction of Simpson’s Paradox had included real rather than hypothetical data. Real data used in the explanation that follows the introduction provide a very good example of the paradox (that is, subgroup trends are different from the aggregate group trend). It is understandable that the final example in that chapter uses hypothetical medical data, but an introductory example could easily be based on real data. A “principle” I find useful is to “present real data, for if you have to manufacture data to give an example of a problem, the problem must be rare and the reader might ignore it.” Students in statistics classes often complain that the manufactured data used in textbook examples are not very realistic.

On page 72 the text under the formula for effect size defines what the symbols for the means are, but ignores the standard deviations in the same formula. This omission might cause confusion for a learner. Also, on page 73 the reader is informed of material on page 159 that will enhance learning. I have found that many students do not take kindly to suggestions to skip ahead to future material, and in this case a better choice would be to explain how this concept differs from what has already been presented on pages 48-49. On those pages the concept of standard deviation was used in conjunction with the normal curve, and the reader might benefit from the recall of previous material.

Given the opportunity to consider the previous interpretation of the standard deviation, the reader could be informed that the effect size is just a measure of distance in standard deviation terms—that is, the standardized distance between the means of the experimental and control groups. The normal curve can then be used to explain the magnitude of that distance. One might call the present writing a missed opportunity for a Vygotsky learning moment (my students use the term scaffolding here). The twins data on page 76 could be used to provide another example of the effect size.

The Brown University example on page 78 is very misleading. The example starts with “Brown University, which could fill two freshman classes just with applicants with SAT verbal scores between 750 and 800 …appeared to give more weight to other factors. For instance, it admitted only one-third of those who scored between 750 and 800.” I think this is supposed to be an example of a school that gives less weight to the SAT; however, given that there were so many applicants who had such high scores, 50% must be rejected just because of lack of room in the freshman class. In addition, if an initial cut-off for applicants might have been set at 650 for a verbal score, that one score might have been heavily weighted for those scoring below 650. Why not explain this case more fully—like the rest of page 78 is devoted to a selection process in algebra?

On page 78 a reference is made to the “Pearson rank-order coefficient” with a comment that “the rank-order coefficient is an approximation of the product-moment correlation”. Actually, the Spearman rank-order coefficient (as it is more frequently called) is the actual product-moment correlation of the ranks rather than an approximation. One should not confuse the readers who were taught these concepts correctly.

In the explanation of the point-biserial coefficient on page 83 the reader might learn “to correlate the chances that a person will get the item right with the person’s total test score.” Now that description might be correct for the estimate of the point biserial given by item response theory applications, but it is not correct for calculating the actual point-biserial coefficient. The correct description would be the correlation between the item score (zero or one) and the person’s total score. Perhaps this situation is an example of where one should be aware of dealing with estimates or scores, and that could be tied with the earlier caution about comparing ranks or rates with scores. Yes, I am picky, but I have learned from Bracey (and others) that I should be careful what I write. And besides, on page 122 he does get it correct, but the reader may not correct that initial misleading definition.

I wonder whether the graphs on pages 98-99 are really stem-and-leaf graphs. That name does not seem to fit what I learned from Tukey’s use of the stem-and-leaf method or what I get when I use a computer program to present data with the stem-and-leaf approach. However, I think the presentation is very clear and the caution presented on page 99 is more than enough to excuse the use of the term. Whatever those graphs are, they are effective. A problem might be that these examples are included in a section titled “Other Ways of Graphing Badly” and the reader might assume there is something wrong with their use.

In the section on “The Nature of Standardized Tests” there is a list of what constitutes a standardized test:

  • the questions are the same for everyone
  • the format of the questions is the same for everyone
  • the instructions given to the students are the same for everyone
  • the time allotted to take the test is the same for everyone
  • the tests are usually administered to a group of students
  • the items on the test have known statistical properties, especially in terms of proportion of test takers who get each item right (pp. 119-120).

Now the above might be Bracey’s wish list, but they do not correspond to what the reader will find in the tests currently described as standardized. With computer-adaptive testing, used for some tests including the GRE, different items are presented to examinees depending on the scoring of items presented previously. As I write this paragraph and come to points 2 and 3, I am looking at an announcement for a conference on Accommodating Students with Disabilities. Now if there are accommodations permissible, are the format, instructions, and time going to be the same for all students? Is the performance of a student who is allowed to type written responses comparable with the performance of one who writes by hand? What if the first student used a word processing program like WORD? Are instructions the same for computer-adaptive and group-administrations of the same test? For the fourth point, why are individually-administered tests not standardized? There is nothing about standardizing tests that differentiates between group and individual administration. The term ‘usually’ may make this point accurate enough, but the idea is misleading. The last point has to do with norms, not standardization. The administration procedure of the test is standardized, but only the data resulting from actual testing can provide information about the statistical properties of the items and the test.

To be fair in commenting on the above list, I will admit that Bracey has addressed some of these points on page 120, but that does not suggest that the list is worth presenting. He presents a great story on pp. 120-121 and relates that story to face validity well, but that does not justify the list. In a book that is intended to encourage one to be a critical reader of testing reports, there is no place for such carelessness.

On page 121 is a list of examples of norm-referenced tests (NRTs) that includes the California Achievement Test and the Comprehensive Tests of Basic Skills but not the TERRA NOVA (TN) that one finds when visiting that company’s web site. A reader may wonder whether TN is omitted because its use in ‘non-standard’ tests such as the Arizona Instrument to Measure Standards makes it something other than a NRT. I give credit in this example for referring to the Stanford Achievement Test as SAT10 to avoid confusion with the SAT. Bracey should have noted that the term NRT is used widely, but actually it is the interpretation, not the test, that is norm-referenced.

On page 122 Bracey does describe ‘norming’ but does not differentiate that activity from standardization. His statement on page 123 that “nationally, 50 percent of the students will always be below average” is suspect given the prevalence of the use of dated norms. But if the scores are reported as stanines, there are only 40% of the norm group scores that are below average. More on this is presented in the book, but not with this principle. It is true that approximately 50% of the scores in the norming group fell below the midpoint called ‘average’, but that is not true of the reported scores (recall the Lake Wobegon effect). That statement with the ‘always’ is more incorrect than principle 18 itself, which might be corrected by using the word “were” rather than “are” if the statistic used as the average were reported correctly.

On page 126 Bracey suggests that local assessments are not as useful as NRTs because they are not backed up by statistics like p values or point biserials. At this point the reader has been told that p represents the level of significance (back on page 71) but has not yet been told what the p value means in this case (which is defined as probability--rather than the actual observed percentage in a group--that got the item correct on page 139). Another problem with this suggestion that national statistics are so important is that the national statistics rarely correspond with local statistics that might represent a better measure for comparison, and the national statistics are very over rated when applied to local curriculum assessments. He might have covered the more important limitations of the national statistics rather than touting their importance. I return to this issue when discussing validity.

A section on Criterion-Referenced Tests (CRTs) follows the section on Construction of a Standardized Test. How misleading is that organization? His opening that NRTs are not much in favor these days and have given way to CRTs is followed by describing how today’s CRTs are nothing like the developers of that type of testing had in mind. This good description is followed by an excellent discussion of setting passing scores and standards-referenced tests. I suggest that his sections on passing scores be disseminated widely. I would hope that in a revision he explains that NRT, CRT, or standards-based tests can be reported with reference to cut scores. This treatment should not be limited to passing scores, as many programs report several levels of cut scores to designate several levels of performers, (e.g., failing, approaching standards, meeting standards, and exceeding the standards).

Now before anyone decides to disseminate the material described in the previous paragraph, why not add a note that William Angoff did not create the procedure called the “Angoff” or probably the “modified Angoff” procedure? Angoff was careful to give credit where credit is due (Wainer, 2006, p.81) and the rest of us should follow that example. Angoff credited Ledyard Tucker. I suspect that neither Angoff nor Tucker would approve of many current test misuses where the methods are mentioned. While on the point of names, coefficient alpha mentioned on page 145 is not “Cronbach’s alpha” (Cronbach, 2004).

On page 145 it should have been mentioned that reliability is a property of the test scores, not of the test (Thompson & Vacha-Haase, 2000). Although obvious to anyone who has considered the effect of restricted range on reliability coefficients, the misconception that reliability is a property of a test is still prevalent in the literature. This book should not perpetuate common misconceptions. And reliability is not restricted to stability, as the treatment on page 145 suggests. The treatment of validity on that page surprisingly does not mention face validity that was introduced on pages 120-121. Any treatment of validity should mention that the data presented by the publisher probably does not represent what will be found in local assessments, and thus validity data are often a property of the local situation rather than the national view addressed by the developer. Principle #27 could be examined using this perspective in addition to the fine treatment already included in this edition of the book.

Normal Curve Equivalents (NCEs) are given little praise in Bracey’s treatment (pp. 157-158), but the same limitations are also found in any other standard scores. After all, NCEs are just standard scores with a mean of 50 and a standard deviation of 21.06 (Bracey might ask, “Who determined the .06 part?)” But IQ or other standard scores have no more theoretical basis than NCEs. NCEs are often derived from national percentile ranks, so they are approximately “normalized” rather than linear-derived standard scores, but most tests come without a description of how the standard scores were derived—so this can be ignored. Rather than dismissing any type of standard score, Bracey could have suggested that these scores are more interval-level measures than are percentile ranks, and thus should always be preferred over percentiles when an average (mean) is calculated. The transformation from NCEs to percentile ranks is as easy as a transformation from any other standard score. Given that NCEs have more score points than other typical reported standard scores, they might be preferred for use in mathematical calculations because they suffer less from assigning students with different raw scores to the same standard score. This last point may be a problem by showing more precision than is warranted; however, I have never read a defense of grouping scores where less misinterpretation of performance compensated for the loss of precision.

On page 160 Bracey reminds the reader that percentages of scores between standard deviation points on the normal curve apply only for normal, bell-shaped curves and not skewed distributions. But this warning confuses the issue of normal curves versus other symmetrical curves that are not skewed. The distinction should be between normal and non-normal, not between normal and skewed. Most observed score distributions that are not skewed are also not normal, frequently being too flat (platykurtic). The reader must know that the percentages between these points do not apply to non-normal distributions even if the distributions are symmetrical.

On page 162 Bracey states that NRTs are designed to be insensitive to instruction. Now the correct statement is that the tests are designed to be insensitive to type of instruction, but not to the effectiveness of instruction. The tests could not be used to measure learning (that is affected by effective instruction) if they were insensitive to the effectiveness of instruction. I agree with most of what he says regarding the misuse of tests for the purpose of evaluating teachers, but used properly tests can be included as one dependent variable in a well-designed experiment to evaluate effective instruction. I will agree that such an experiment is rare, indeed.

The major section of the book is entitled “Testing: A Major Source of Data—and Maybe Child Abuse”. He gives so many examples of improper interpretation of test results that I cannot argue with the title. I just hope that in the future editions, and perhaps in his future writing in general, he’ll focus a bit on these comments rather than just perpetuate some of the misunderstandings that are prevalent in the literature today. I did not include every suggestion I have for improving the book; thus I do not want to hear from anyone about what I missed. However, I would respond to comments about my being incorrect about topics where I might have cited authorities rather than just pretending to be an expert.

And Jerry, whether you use any of this or choose not to, I will continue to look forward to reading your views on educational issues of the world.

References

Cronbach, L. J. (2004). My current thoughts on coefficient alpha and successor procedures. Educational and Psychological Measurement, 64, 391-418.

Wainer, H. (2006). Book review: Phelps, Richard P. (Ed), Defending Standardized Testing. Journal of Educational Measurement, 43, 77-84.

Thompson, B. & Vacha-Haase, T. (2000). Psychometrics is datametrics: The test is not reliable. Educational and Psychological Measurement, 60, 174-195.

About the Reviewer

Darrell Sabers is Professor in the Department of Educational Psychology at the University of Arizona. His research specialty is applied psychometrics, especially focused on educational testing and research. Darrell teaches courses in educational measurement (introduction and advanced theory) and in research methods. His current research projects include test analysis for a test of early mathematics achievement, scale construction and validation in deaf education, and validity studies in intellectual assessment.

Copyright is retained by the first or sole author, who grants right of first publication to the Education Review.

Tuesday, March 18, 2025

Willingham, D. T. (2017). The reading mind: A cognitive approach to understanding how the mind reads. Reviewed by Caroline Whitley, University of Mississippi

Education Review-a journal of book reviews

Willingham, D. T. (2017). The reading mind: A cognitive approach to understanding how the mind reads. Jossey-Bass.

256 pp.                   ISBN: 978-1119301370

Reviewed by Caroline Whitley
University of Mississippi
USA

Daniel Willingham's The Reading Mind: A Cognitive Approach to Understanding How the Mind Reads offers an in-depth look at the cognitive science behind reading. Willingham, a cognitive psychologist, breaks down the complex processes that happen in the brain when we read, from letter recognition to comprehension and beyond. His writing is engaging and backed by research, making this book a useful resource for educators looking to better understand how students develop reading skills.

Willingham structures the book to guide readers through different aspects of reading, explaining how the mind processes words, sentences, and entire texts. He begins with the idea that reading is more than just recognizing words. It involves building meaning from sentences, connecting ideas, and applying background knowledge. The introduction, where he discusses what he calls the "Chicken Milanese Problem," sets the tone for the book. He uses this example to show that people often engage with text in a way that feels familiar rather than deeply understanding it, which is a key issue in reading comprehension.

One of the most compelling parts of the book is the chapter on reading comprehension. Willingham explains how readers do not just absorb words on a page. They actively construct meaning by connecting information across sentences and paragraphs. He introduces the concept of a "situation model," which is essentially the mental representation a reader builds while reading. Skilled readers do this naturally, while struggling readers may fail to make necessary connections, leading to gaps in understanding. A particularly eye-opening discussion is how background knowledge impacts comprehension. Research shows that students who already have some knowledge about a topic tend to understand related texts better, even if their general reading skills are not as strong. This challenges the idea that comprehension skills alone determine reading success and highlights the importance of building students’ knowledge base across different subjects.

Another interesting section of the book focuses on how readers interpret meaning using grammar, prior knowledge, and inference making. Willingham provides examples of how simple sentences can be misleading if a reader lacks context. He also discusses "idea webs," the mental structures readers use to connect information within a text. His explanations make it clear that reading comprehension is not a passive process. It requires active engagement, logical reasoning, and a foundation of background knowledge to be effective.

Of all of the arguments Willingham made in the book, one of the strongest is that teaching reading strategies is not enough on its own. While strategies like summarizing, questioning, and making predictions can help students, they will not be very effective if students do not have the background knowledge to understand what they are reading in the first place. He emphasizes the need for a curriculum that builds students’ knowledge across various subjects rather than focusing solely on teaching generic reading skills. This perspective is especially relevant for educators who want to improve literacy instruction in meaningful ways.

The Reading Mind is an insightful and practical book that bridges cognitive science and classroom instruction. Willingham explains complex ideas in a way that is easy to follow while still being research driven. Some sections can be a bit heavy on studies and theories, but they are balanced with real world applications that make the book useful for teachers. His argument that comprehension is deeply tied to knowledge rather than just reading skills alone is a valuable takeaway for anyone involved in education. This book is a must read for teachers who want to better understand how their students think as they read and how they can improve reading instruction in the class.

Education Review is supported by the Mary Lou Fulton College for Teaching and Learning Innovation, Arizona State University. Copyright is retained by the first or sole author, who grants right of first publication to the Education Review.

Hess, Frederick M. (ed.) (2008). <cite>When Research Matters: How Scholarship Influences Education Policy</cite>. Reviewed by Mark Oromaner

Hess, Frederick M. (ed.) (2008). When Research Matters: How Scholarship Influences Education Policy . Cambridge, MA: Harvard Edu...