back

by NaOH·14y ago·view on hn ↗
I worked for a publisher of standardized tests for a number of years, though this company did not develop the assessments used by the State of New York. What this teacher doesn't seem to appreciate is the process behind how these tests are written. They are, seemingly by the nature of this society, produced in a manner which can't capture distinctive teaching styles like this woman has.

The process of making a test can start with whatever a state legislature has mandated will be assessed. Mind you, before that, there is all the negotiations and politicking that takes place. Surprising to most, this involves state education leaders, business people, religious folks, politicians, parents, etc. It's a kitchen sink of divergent interests with everyone claiming to have the best interests of the children at the fore.

Once the legislation is in place and a contractor has been secured to aid with development, there are loads more meetings and committees deciding what is appropriate assessment within each subject (math, reading, writing, etc.). Again, there are loads of different people with loads of different interests, all of whom believe they are thinking foremost about the children.

It's also in this stage that whatever research or trends in assessment styles will be considered (though some takes place earlier, too). There's usually a mentality of "You go first" to new assessment techniques. States are more willing to try something if another state has already done something similar and there is publicly available data to support the perceived efficacy.

At the next stage is actual development of the assessment materials. We would split this up, part of it being done in-house and a bunch contracted to teachers around the state. Yes, we tried to get teachers from every district. For the contracted work, this would mean paying teachers to write a number of questions for a specific test (say, fifth-grade math). These people got paid for each question they wrote regardless of the quality or usability of what they submitted.

The worst material to write was probably math simply due to the dry nature of the subject and the fact that creative approaches to math are usually verboten in education here. Reading tests were often the most difficult to develop. The work on the tests wasn't bad, but securing copyright permissions and, often, permission to edit was brutal. If there was a magazine piece, the complications were often much worse because usage rights might have to be secured from multiple parties (publisher, author and photographers).

Mind you, this was also in the late '90s, so the Internet wasn't as useful a tool for tracking down rights holders or potential materials, and email was still a secondary means of communication, definitely behind the phone and often behind the fax, too. Securing rights for all the materials we wanted to use took months just because of how hard it was to find people and communicate with them. And states didn't have much of any budget to pay, so securing rights at minimal cost was a big hurdle. Often, the best pieces were never used due to how much a rights holder wanted.

So questions would come in from all over the state, then we would clean them up. That was multi-layered work. It might mean simple grammar and punctuation fixes, but it also meant correcting the format mandated by the state education departments. For example, when I was doing this work, states would not allow us to put a negative in the question. But there were loads of these types of rules, like making certain there was parallel structure among answer choices, not having any choices significant;y shorter or longer, etc.

Once we had done an initial tightening of the new bank of material, all the teachers we'd contracted and state administrators were brought in for a week of refinement and further development of materials. These were simultaneously productive and political sessions. A lot of work would be done, but there was also a lot of on-site jockeying. Teachers would say things like, "This is a great story, but my kids won't be able to relate to it." State administrators would hear this a few times about a piece and then pull the material from any further consideration, not even pilot testing. Quality was often a secondary consideration to how teachers felt their students would do, and it was sometimes tertiary to other teacher goals (what they believed was important, their personal agendas, etc.).

Those last few steps would then repeat themselves. We would tighten up the work that had been developed, the graphics department would develop accompanying graphics where needed and handle page layout, proofreading was a persistent process, and then we would bring the teachers back in for another review of the nearly final materials.

Then we would do another round of tightening-up the material. Proofing, requesting minuscule tweaks from the graphics department, getting state approval for any substantive change (no matter how minor), etc. That was when we could begin building an actual test using these new materials and existing questions from previous tests. We'd also begin to development the accompanying manuals which instructed the schools how to handle the materials and the teachers how to administer the tests. As you can imagine, these had to be perfect. When you have a 100-page document with loads of instructions around specific details, errors are not permissible.

(I haven't done this work in 12 years, but to this day my eyes proofread nearly everything that comes before them. I can be at a simple restaurant, and the menu will list "pan fried chicken." I instinctively note the missing hyphen.)

Of course, all that development work only went to pilot materials. I don't remember exactly, but a student might take a test that was about 70% questions that counted and the rest were new questions being evaluated. Once the tests came back, data analysis was run on everything, enabling us to see what worked and what didn't. Sometimes a question was too hard or too easy, sometimes one group of people simply had issues with a question. My memory is hazy, but I want to say that about one-third of the questions that were piloted became usable. Maybe 10-20% of them got re-piloted because the data showed a way we could possibly fix the question (e.g., one of the answer choices was too attractive, so a re-write of that might be enough of a fix to make the question worth trying again).

On the other side was the scoring for written questions. We had the state-issued rubrics, and those were our guiding force. I (and others) would train the part-time people we hired to do this scoring. The company I worked for hired these people largely off of a standard bank of psychological assessments. The company owner felt these gave all the information we needed to evaluate these potential employees.

Easily, the biggest challenge was getting scorers to accept the rubrics. A student might write a quality piece about something, but it might have been well off topic or not sufficiently on topic based on what the state wanted to assess. During training sessions, I spent a lot of my time diffusing anger from these people and getting them to focus on the rubrics. Gently humor was key in that regard, and I don't recall anyone proving to be a long-term problem in terms of accepting the rubrics.

The other big challenge to this work was the repetition. Reading the answers to the same questions over and over was mentally challenging for people. I don't blame them. Most kids of a specific age aren't too creative when fed a question for a state test. For example, ask them who is a public figure they admire and why, and you're likely to get the bulk of the answers focusing on just a few people (athletes, popular music stars, etc.). For the written assessments, 10% of the student materials were scored twice (by separate people) to ensure accuracy of grades and as a way to identify issues with potential scorers.

I've tried to refrain from too much commentary, but there is no doubt that the materials developed for tests are beaten down throughout the process by bureaucracy and various interests. It's much like the corporate world when the firm has way too many meetings in the course of developing something and there is a leadership vacuum. Oh, sure, there is a person or two who is technically leading things and may have veto power, but there are far too many diverse interests for anything of distinct quality to emerge.

With one of the states for which my employer did work, a woman like the teacher in the link would be invited to participate in the following year's development. The lead state administrator always referred to this as "getting that person's buy-in." And truthfully, it seemed to work because the teachers brought in for this reason felt like they had a voice in the process. None that I saw seemed to appreciate the depth of the whole process, so they all seemed to think they had made a difference in the development of the tests.

More specific to the author of the link, she seems like she's probably a good teacher, better than most. The education system, especially when it comes to statewide assessments, isn't prepared to appreciably handle outliers like her. She probably knows that. She probably knows the real battle to change these kinds of tests is not one she's prepared to tackle. I don't blame her. Nor do I blame her for making a public critique like she did.

8 comments
“I told you, didn’t I, about hearing Noam Chomsky speak recently? When the great man was asked about the chaos in public education, he responded quickly, decisively, and to the point: “Public education in this country is under attack.” The words, though chilling, comforted me in a weird way. I’d been feeling, the past few years of my 30-plus-year tenure in public education, that there was something or somebody out there, a power of a sort, that doesn’t really want you kids to be educated. I felt a force that wants you ignorant and pliable, and that needs you able to fill in the boxes and follow instructions. Now I’m sure.”

This speaks to a rather odd state of mind on the part of the teacher. “The great man”?

Believing that the state of education and testing is the result of some powerful villain behind the curtain requires more credulity than simply attributing it to the massive complexity and inertia you describe.

This speaks to a rather odd state of mind on the part of the teacher. “The great man”?

Believing that the state of education and testing is the result of some powerful villain behind the curtain requires more credulity than simply attributing it to the massive complexity and inertia you describe.

Certainly Chomsky is a great man. He's the founder of modern linguistics. He wrote the paper that demolished behavioral psychology, paving the way for cognitive psychology. He invented the hierarchy of formal languages that is still taught in every Computer Science program in the world. He is the most cited academic in the world. He is the most influential thinker in two completely unrelated fields: linguistics and political activism. He's an "Institute Professor" at MIT, which is the highest honor that MIT bestows. (I think that there are fewer of them at MIT than there are Nobel Laureates.)

Just what is there not to consider great?

As to "some powerful villain", where did either Chomsky or the OP ever assert that the "villain" isn't a system? Perhaps a system comprised largely of complexity and inertia?

I know for a fact that Chomsky believes something along these lines. He's stated many times that the harmful influences that he describes are not the result of a cabal, as he is often misattributed as asserting, but that he is and always has been speaking about the effects of a complex system, whose intent is not the intent of any individual part of the system.

You make the case much more convincingly than the OP. The passage I quoted came across to me as uncritical fawning.
As opposed to your uncritical lack of appreciation?
"Never ascribe to malice that which is adequately explained by incompetence."
I read that quote too often. Greed exists and often masquerades as incompetence. Do not be always that adamant.
I found it jarring, too.

Fun link: a fake interview with Chomsky: http://proteinwisdom.com/?p=16470

Warning, link is to a right wing propaganda website.
Do you give equal time fair warnings regarding left-wing propaganda websites? Hmmm?
>Public education in this country is under attack

I think it would be far more accurate to say that our entire culture is under attack, from multiple vectors, to reduce the current upwardly ambitious mass populous to neo-feudalism.

Keeping the peasants stupid is just one tactic in a complex strategy.

what is odd about admiring noam chomsky?
Nothing. I admire him. But I wouldn't cite him in such a hero-worshiping manner. Doing so would detract from my argument and undermine my credibility.
Why does a person have to admire?

(FYI writing the first letters of a person's first and last name in lower-case is not very polite)

And what kind of middle school manages to get Noam Chomsky invited for a talk? Is that something he does normally?
Doesn't sound like he did. Sounds like the teacher heard him speak, and then told students about it.
Thanks for your fascinating comment.

Unfortunately, all the excuses in the world don't change the fact that these tests hurt education and hurt children. The production of these things is a shameful act.

I thought the tests we had in the UK were bad, and they are, but the US obviously has a far greater problem. I hope this changes, because you are hurting your children and damaging your future.

>Unfortunately, all the excuses in the world don't change the fact that these tests hurt education and hurt children.

The rub is that: so do terrible teachers.

It's inevitable in a world where bureaucrats take a top down approach to test taking and measurement that you end up with people who "teach the test". It sucks.

But, when left to their own devices, some teachers don't teach at all. Teachers follow the same distribution of personalities as any other occupation on this planet; some are good, some are bad, most are average. These standardized tests were an attempt at saying, "You MUST teach at least this well", with "well" being defined by factoids. So, on one hand you have a free-for-all, where great teachers influence and inspire, while poor teachers ruin lives. The other side is a middle ground where what is to be taught is quantified, and everyone gets a mediocre (at best) education.

What we need a hybrid system that allows good teachers to do what they want, while bad teachers are held accountable and forced to, at the least, teach the test.

But, when left to their own devices, some teachers don't teach at all. Teachers follow the same distribution of personalities as any other occupation on this planet; some are good, some are bad, most are average.

In my experience there are precious few teachers who don't teach at all. In fact, I never had a single one that bad. At worst, they hand you the text book and make you read it over the course of the year and give you some tests on it.

The best teachers don't teach much really. Rather they inspire. The worst teachers may teach a lot, but they make you despise the material.

It seems to me that having teachers be forced to teach to standardized tests will kill off the first type of teacher, while having little or no effect on the second kind of teacher.

The worst of all worlds.

> The best teachers don't teach much really. Rather they inspire. The worst teachers may teach a lot, but they make you despise the material.

Going to my 11th grade Biology class every day, I used to joke with a couple friends "Do you think we'll be learning today?", "Ha ha, Probably not." The class was fun, not very serious, with the bulk of the period spent listening to our teacher tell stories. All that and the homework was always easy.

But at the end of the year, while studying for the final, we realized that we had learned a huge amount of biology. Our teacher had spent the classes telling us stories that related to the material, and the homework was easy because she inspired us to be interested in the questions being asked. I didn't remember toiling over the Krebs[1] and Calvin[2] cycles, but I certainly knew them. More importantly though, I found them interesting.

[1]http://en.wikipedia.org/wiki/Citric_acid_cycle

[2]http://en.wikipedia.org/wiki/Calvin_cycle

> What we need a hybrid system that allows good teachers to do what they want, while bad teachers are held accountable and forced to, at the least, teach the test.

Bad teachers being held accountable is a good idea. It is such a good idea that bad teachers should lose their jobs. In some cases they should even lose their certification.

Bad teachers teaching bad tests is worse than bad teachers left to their own devices.

"Factoids" is a laughable metric for learning. (Especially when you're looking for factoids in a long-form response. If you want factoids, you need to ask for factoids. This test asked for complex thought, the grading should require complex thought.)

>"Factoids" is a laughable metric for learning.

Not if learning those factoids allows one to move up the chain, from what is probably a terrible teacher to one that is hopefully, better. Or, those factoids might get you a job where one can begin to learn outside of a scholastic setting which isn't ideal for every personality. I realize it's not ideal. That's why it's such a difficult problem.

I come from a country (and province) with an incredibly powerful teacher's union. So I've seen what happens when teacher's are not held accountable for anything.

>"This test asked for complex thought, the grading should require complex thought"

Do you trust the average teacher to be a good judge of "complex thought"?

Bad teachers teaching bad tests is worse than bad teachers left to their own devices.

How is it possible that a student is worse off having the knowledge necessary to pass a test than to have no knowledge?

This test asked for complex thought, the grading should require complex thought

Kids are dumb. Their "complex thought" is really quite simple. Further, the iPhone required a stunning amount of complex thought to create, yet I can asses its functionality with only simple observation. I doesn't necessarily require complex thought to asses the product of complex thought, particularly when that's complex thought at grade level X.

Unfortunately, all the excuses in the world don't change the fact that these tests hurt education and hurt children.

I agree, but my goal was more to share what I've seen of the process than to judge it. I think folks here are percipient enough to make those judgments independently. On top of that, I've long been out of the field and am far too ignorant too suggest alternatives, and that's what we seem to need more than critiques.

  I've long been out of the field
Speaking as someone that came out of high school a couple years ago and watched No Child Left Behind start taking effect, I've seen two different styles of test: one from the mid-90's, when it sounds like you were involved, and one from after No Child Left Behind. The NCLB tests today are utterly unlike the tests that your methods produced. Basically, instead of trying to rank students against each other, they began ranking every student against the same fixed hurdles. On top of that, they started testing once every year or two rather than once every four or five years, and question quality dropped precipitously as they just ran out of material. The changes were, to me, obvious, and the effects were disastrous. Null-value questions ("What shape is the end of a pencil? Cone, square, pyramid.", on a 7th-grade test), bad and wrong essays (as described in the OP), science questions that asked for rote memorization rather than actual understanding and thinking, and so on.

These observations, though amateur and unscientific, suggest a few solutions. Basically, go back to relative rankings, do a lot less testing, and start rewarding uniqueness. Teachers need to have time to actually teach outside hitting every checkbox on list for the end-of-the-year test. Teachers need to be able to teach to their own students, rather than to the lowest-common-denominator child in rural Alabama. Teachers need to be able to encourange and guide every student individually, developing artists, scientists, thinkers, and doers, rather than having to force every single child into the same standardized mold.

</rant>

Right, I was in the field from something like 1995-2001, so I was before the No Child Left Behind style of testing. And I've been nowhere near the field in my work since.

If I were to guess, the types of questions you received were different, but the general process of how the tests were made remained the same. And that's the bureaucratic, lowest-common denominator mentality that pervades broadly given tests. And I agree with you that "[t]eachers need to be able to teach to their own students." I think that was a big part of the OP's point since that's something she does.

It's all part of the conundrum: Everyone wants good teachers and accountability in education, but there doesn't appear to be a high-level way of making such assessments that can pass the necessary political muster.

I meant it when I thanked you for the comment. It was a fascinating insight into how these tests actually come about. Thanks for writing it!
Unfortunately, all the excuses in the world don't change the fact that these tests hurt education and hurt children.

They also help, by identifying bad teachers and bad schools.

Do you have any data suggesting they hurt more than they help?

In the UK, we have spent ages and a lot of effort in testing, league tables, improvement plans, improvement targets for schools.

There is a slowly growing realisation that the result may not be good for the children.

See the Wolf report

https://www.education.gov.uk/publications/standard/publicati...

and also a short OFSTED report into Maths teaching, see the newspaper summary below. The full report may be of interest to you as it explains the pedagogy of mathematics well.

http://www.telegraph.co.uk/education/2982483/Ofsted-testing-...

Teaching is a highly dimensional task. Assuming you could devise a metric space adequate to the task, my teaching at any point would be represented by a set of coordinates. The norm of the coordinates of my 'point' in the space might be higher or lower than the norm of another teacher. How do you decide which one of us is less 'bad'?

Your first link, near as I can tell, doesn't even attempt to address the question of whether standardized tests are beneficial or harmful. It seems to be about the merits of vocational vs higher education.

Your second link provides no data on whether testing is good for children. It merely shows that actual teaching methods do not conform to what the author's believe are the best teaching methods. No data is provided on student outcomes.

Teaching is a highly dimensional task...How do you decide which one of us is less 'bad'?

Any goal-oriented system is designed to maximize some arbitrary objective function. With standardized tests, you are forced to write down your objective function and admit it's an arbitrary choice.

What benefit do we receive from having an unspecified objective function and no uniform method of measurement?

> They also help, by identifying bad teachers and bad schools.

Aside from the arguments posted in sibling comments about what's being measured, you also have a problem of unaligned incentives. In particular, your claim only has a chance of being true if the students are actually interested in scoring as high as they can on the exam. Even at the AP level (I teach at university level and interact with secondary teachers that teach smart, high-level students taking CS), there are students who have decided that they don't care about the subject, or maybe even like the subject but have no (perceived) benefit from a high score due to their chosen college not giving credit in that subject or whatever. Such students may leave their answer book blank, or doodle in it, or maybe just blast through for the easy points and finish early and not worry about thinking about it.

This isn't even necessarily a particularly irrational choice on their part!

But it's a strong argument why the exams shouldn't be used to evaluate the teacher or the school. In a lot of places the students aren't permitted to opt out of the exam, even if they don't care about it, but there's no penalty to the student for taking a dive on it (and any penalty you could try to assess would have false positives and false negatives and still not motivate many of the students with differently-aligned incentives).

All of these problems are going to be a million times worse on a general-education primary- or secondary-level assessment than they are on AP exams.

The data has proved to be virtually useless in assessing teachers. See http://garyrubinstein.teachforus.org/2012/02/26/analyzing-re...

They're useful for showing students are underperforming. But you can't say anything about teacher quality. In fact I suggest if you took the teachers from the "worst schools" and put them in the "best schools" and vice-versa the results the following year wouldn't look noticeably different.

Does that mean that teacher have no impact at all. I don't believe that. But I do think the large gaps in achievement are systemic problems more than the problem of teacher quality at certain schools.

The fact that some teacher can't see the correlations using an inappropriate plotting technique (he uses a scatterplot, he should use hexbin or other density plot) doesn't mean they aren't there. And the numbers he gathers and dismisses show the correlation actually does exist.

A simple numerical example, where the correlation is guaranteed (i.e., I took y=x+noise):

http://i.imgur.com/rrmUI.png https://gist.github.com/2183927

Most of the correlations he expects to find are present. They are noisy, but present.

The fact that in one case, reality is "contrary to what every teacher in the world knows" just suggests maybe teachers don't have a great grip on reality.

...if you took the teachers from the "worst schools" and put them in the "best schools" and vice-versa the results the following year wouldn't look noticeably different...Does that mean that teacher have no impact at all?

Not quite, but almost. It means that variation among existing teachers impact less than the size of measurement, and the current crop of teachers are basically interchangeable cogs.

> They also help, by identifying bad teachers and bad schools.

Not really. Here in Chicago, the Noble charter high schools will often teach directly to the ACT. This results in some dramatic boosting of ACT scores, with some schools taking kids from the dysfunctional CPS System and getting a school average of 23. Unfortunately, once they reach college, they see little academic success despite being straight-A students with 25+ ACT scores.

This has a detrimental impact on the curriculum:

1) The English program structured heavily around basic reading comprehension, with little to no emphasis on writing composition. A students understanding of essay composition s roughly: "Organize the things you want to talk about into paragraphs... then write a conclusion." However, to their credit, they're really good at reading test questions.

2) Math is focused on teaching Pre-algebra fundamentals and then layering on test-specific Algebra, Geometric (with that goofy proof system), and basic Trig. It's a sad, narrow sample of our already sad & narrow HS math curriculum. It covers few "advanced" Algebra and Trig subjects. This means anyone who has to take "college math" will need courses in Trig and Precalc in college, with the possibility of an Algebra refresher course before proceeding.

3) Social Studies and Science? All rote-to-test. Students are drilled on step-by-step procedures on how to interpret graphs that'll score correctly on tests... without giving them actual knowledge on how to critically think about information - be it historical or scientific.

NCLB schools do the same song-and-dance, except with a much less rigorous test. If you've actually seen the questions on most NCLB tests, you'd be disappointed. Unfortunately, the composition of such tests is so political and messy it's impossible to provide any measure of quality.

You cannot assume testing will provide you accurate information or a better outcome for students. If you're going to implement a testing regiment, you need to be very mindful of the Observer Effect: You can very easily change the outcome by measuring it. This is not an easy problem to solve.

Your issue is with the objectives of the school system, not testing. Tests measure whether a school meets it's objectives, they don't define the objectives.

Similarly, if your manager writes a bunch of nonsensical unit tests, it's ridiculous to blame unit testing if you wind up building the wrong product.

Annual tests end up directly costing students a month of education time, I think you need to prove benefit and not the other way around. As to their value, collages still trust SAT's and GPA more than any of the state tests in large part because they are uniformly terrible. If you really want to test teachers then randomly assign 1-2 tests to each student, it's just as statistically valid and takes ~1/3th the time.
But that's the thing, they don't identify bad teachers. And they do penalize good teachers. Being taught to pass these tests is not the same as being educated. All the tests do is identify teachers who don't put the majority of their teaching effort into getting children to pass the tests, to the detriment of everything that matters.
Do they identify bad schools? Or do they identify schools with children that are harder to teach?

This is a genuine question.

I'm also a bit disappointed about the lack of rigorous evidence base with education. (Well, with everything, really.) In the UK the Department for Families and Schools has a lot of power over education. In theory that should help with an evidence base because they should be flowing good research down through to schools. But there's such a political atmosphere about teaching that interference from government is often seen as unwelcome. And, sometimes, that attitude is correct be government is suggesting something that's stupid.

There's also a problem that bad teachers are difficult to remove from teaching.

The other big challenge to this work was the repetition. Reading the answers to the same questions over and over was mentally challenging for people. I don't blame them. Most kids of a specific age aren't too creative when fed a question for a state test. For example, ask them who is a public figure they admire and why, and you're likely to get the bulk of the answers focusing on just a few people (athletes, popular music stars, etc.). For the written assessments, 10% of the student materials were scored twice (by separate people) to ensure accuracy of grades and as a way to identify issues with potential scorers.

I'm curious - do you happen to know whether the essay portions ever offered any actual statistical utility above and beyond the multiple choice ones? As in, do they actually measure anything that the multiple choice questions can't?

I spent several years teaching SAT and GRE classes, I always noticed that I could predict people's essay scores pretty accurately based on their multiple choice scores. Percentile ranks always tended to be very close between the sections, which always made me wonder whether there was any point in having the essay at all. I always suspected that since the overall level of competence was so low on the SAT (sadly, the GRE was not all that much better...), anyone that was writing at the level where actual quality of writing matters was already scoring top marks on the essay, so they were essentially unmeasured by the scale.

I realize that when it comes to setting up these tests, "include an essay" is likely a political mandate, because people for some reason think that essay sections are "more fair" to "bad testers" or "fluid thinkers" or something like that (I'm pretty sure ETS was more or less forced to include an essay for this reason), so it probably wasn't even an option not to include them. But I'd be curious to know whether anyone ever looked into whether they actually told you anything you didn't already know. ETS does not provide this data or analysis, otherwise I'd check it myself as it relates to the SAT.

I'm curious - do you happen to know whether the essay portions ever offered any actual statistical utility above and beyond the multiple choice ones? As in, do they actually measure anything that the multiple choice questions can't?

I believe written assessments were relatively new to standardized tests when I was in the field. My understanding was that the purpose was to assess writing, just like we do with, say, math or reading comprehension. Multiple-choice questions can't evaluate a student's ability to respond to an open-ended question, convey ideas, stay on topic, demonstrate grammar and punctuation, etc.

But I didn't get the impression that adding writing components to the test added anything of practical value. Sure, it's one more score-based judgment that can be assigned to each student and be used to assess schools and districts, but as far as I could tell it wasn't benefitting student education. And it certainly came with a high price tag. But I was definitely not in classrooms, so that's a key caveat to anything I share.

Does writing seem any better? Well, these written assessments have been widespread for about 20 years or so. I don't get the sense from what I see that people nowadays are generally good writers. But my exposure is certainly limited, even if others I speak to seem to agree, and I'm not certain if enough time has elapsed whereby we would broadly see the effects of such a change in educational assessment.

I'm curious - do you happen to know whether the essay portions ever offered any actual statistical utility above and beyond the multiple choice ones?

I don't think statistical utility was the goal - the goal might have just been to reduce math from 50% of the test to 33% of the test.

The (unproven, but widely repeated) story is that after Prop 209, too many of the wrong type of people were getting into UC due in part to high math scores [1]. California proposed dropping the SAT requirement as a result, and the College Board came up with a way to reduce it's demographic impact (and keep CA students paying them).

[1] http://professionals.collegeboard.com/data-reports-research/...

That the process is onerous does not make it correct.

That a good teacher is an exceptional case that we simply cannot handle with our bureaucracy is a thought that should terrify you to your core.

We get the quality we demand. Stop apologizing for the inadequate status quo.

I liked the middle paragraph in your comment. The rest was pointless garbage. The GP gave a very thorough and enlightening description of the test development process. There's no sense in attacking the messenger: he wasn't positioning himself as an apologist; he was merely informing us what goes into these tests.
Fair enough. I didn't think of my post as an attack, but if it read as one, I should consider softening my language more.

But from my perspective, the subtext of the post was "you don't know how hard this is, you should just be happy about what we do have". I think we need to demand more, and those of us who can contribute to educational reform, contribute more.

I think a problem is that humane teachers are made to be "outliers". She correctly perceives these standardized tests as an attack against students and decent teachers. From bureaucracies like the ones you describe. If these bureaucracies were more efficient, I don't think that'd improve the situation; probably even strengthen their attacks.
"worst material to write was probably math simply due to the dry nature of the subject"

Math is one the most fascinating, creative, and empowering pursuits of our species.

Good teaching consists of intelligently applying basic principles to particular situations, which are themselves understood by subjective criteria. It isn't hard to differentiate a good teacher from a bad one, once you take time to look, though it does sometimes take some imagination and observation.

The tests are efforts to identify good teaching on a general and objective basis. That's probably doomed to fail from the start, and what theoretical potential is available is quickly crushed by the political components of the process.

The real question is why we need general and objective measures in the first place. The answers have nothing to do with education, and everything to do with the governance of educational institutions. That governance is the real problem, and nothing's going to matter until it's addressed.

"Creative approaches to maths are usually verboten in education"

I will quote this.