The process of making a test can start with whatever a state legislature has mandated will be assessed. Mind you, before that, there is all the negotiations and politicking that takes place. Surprising to most, this involves state education leaders, business people, religious folks, politicians, parents, etc. It's a kitchen sink of divergent interests with everyone claiming to have the best interests of the children at the fore.
Once the legislation is in place and a contractor has been secured to aid with development, there are loads more meetings and committees deciding what is appropriate assessment within each subject (math, reading, writing, etc.). Again, there are loads of different people with loads of different interests, all of whom believe they are thinking foremost about the children.
It's also in this stage that whatever research or trends in assessment styles will be considered (though some takes place earlier, too). There's usually a mentality of "You go first" to new assessment techniques. States are more willing to try something if another state has already done something similar and there is publicly available data to support the perceived efficacy.
At the next stage is actual development of the assessment materials. We would split this up, part of it being done in-house and a bunch contracted to teachers around the state. Yes, we tried to get teachers from every district. For the contracted work, this would mean paying teachers to write a number of questions for a specific test (say, fifth-grade math). These people got paid for each question they wrote regardless of the quality or usability of what they submitted.
The worst material to write was probably math simply due to the dry nature of the subject and the fact that creative approaches to math are usually verboten in education here. Reading tests were often the most difficult to develop. The work on the tests wasn't bad, but securing copyright permissions and, often, permission to edit was brutal. If there was a magazine piece, the complications were often much worse because usage rights might have to be secured from multiple parties (publisher, author and photographers).
Mind you, this was also in the late '90s, so the Internet wasn't as useful a tool for tracking down rights holders or potential materials, and email was still a secondary means of communication, definitely behind the phone and often behind the fax, too. Securing rights for all the materials we wanted to use took months just because of how hard it was to find people and communicate with them. And states didn't have much of any budget to pay, so securing rights at minimal cost was a big hurdle. Often, the best pieces were never used due to how much a rights holder wanted.
So questions would come in from all over the state, then we would clean them up. That was multi-layered work. It might mean simple grammar and punctuation fixes, but it also meant correcting the format mandated by the state education departments. For example, when I was doing this work, states would not allow us to put a negative in the question. But there were loads of these types of rules, like making certain there was parallel structure among answer choices, not having any choices significant;y shorter or longer, etc.
Once we had done an initial tightening of the new bank of material, all the teachers we'd contracted and state administrators were brought in for a week of refinement and further development of materials. These were simultaneously productive and political sessions. A lot of work would be done, but there was also a lot of on-site jockeying. Teachers would say things like, "This is a great story, but my kids won't be able to relate to it." State administrators would hear this a few times about a piece and then pull the material from any further consideration, not even pilot testing. Quality was often a secondary consideration to how teachers felt their students would do, and it was sometimes tertiary to other teacher goals (what they believed was important, their personal agendas, etc.).
Those last few steps would then repeat themselves. We would tighten up the work that had been developed, the graphics department would develop accompanying graphics where needed and handle page layout, proofreading was a persistent process, and then we would bring the teachers back in for another review of the nearly final materials.
Then we would do another round of tightening-up the material. Proofing, requesting minuscule tweaks from the graphics department, getting state approval for any substantive change (no matter how minor), etc. That was when we could begin building an actual test using these new materials and existing questions from previous tests. We'd also begin to development the accompanying manuals which instructed the schools how to handle the materials and the teachers how to administer the tests. As you can imagine, these had to be perfect. When you have a 100-page document with loads of instructions around specific details, errors are not permissible.
(I haven't done this work in 12 years, but to this day my eyes proofread nearly everything that comes before them. I can be at a simple restaurant, and the menu will list "pan fried chicken." I instinctively note the missing hyphen.)
Of course, all that development work only went to pilot materials. I don't remember exactly, but a student might take a test that was about 70% questions that counted and the rest were new questions being evaluated. Once the tests came back, data analysis was run on everything, enabling us to see what worked and what didn't. Sometimes a question was too hard or too easy, sometimes one group of people simply had issues with a question. My memory is hazy, but I want to say that about one-third of the questions that were piloted became usable. Maybe 10-20% of them got re-piloted because the data showed a way we could possibly fix the question (e.g., one of the answer choices was too attractive, so a re-write of that might be enough of a fix to make the question worth trying again).
On the other side was the scoring for written questions. We had the state-issued rubrics, and those were our guiding force. I (and others) would train the part-time people we hired to do this scoring. The company I worked for hired these people largely off of a standard bank of psychological assessments. The company owner felt these gave all the information we needed to evaluate these potential employees.
Easily, the biggest challenge was getting scorers to accept the rubrics. A student might write a quality piece about something, but it might have been well off topic or not sufficiently on topic based on what the state wanted to assess. During training sessions, I spent a lot of my time diffusing anger from these people and getting them to focus on the rubrics. Gently humor was key in that regard, and I don't recall anyone proving to be a long-term problem in terms of accepting the rubrics.
The other big challenge to this work was the repetition. Reading the answers to the same questions over and over was mentally challenging for people. I don't blame them. Most kids of a specific age aren't too creative when fed a question for a state test. For example, ask them who is a public figure they admire and why, and you're likely to get the bulk of the answers focusing on just a few people (athletes, popular music stars, etc.). For the written assessments, 10% of the student materials were scored twice (by separate people) to ensure accuracy of grades and as a way to identify issues with potential scorers.
I've tried to refrain from too much commentary, but there is no doubt that the materials developed for tests are beaten down throughout the process by bureaucracy and various interests. It's much like the corporate world when the firm has way too many meetings in the course of developing something and there is a leadership vacuum. Oh, sure, there is a person or two who is technically leading things and may have veto power, but there are far too many diverse interests for anything of distinct quality to emerge.
With one of the states for which my employer did work, a woman like the teacher in the link would be invited to participate in the following year's development. The lead state administrator always referred to this as "getting that person's buy-in." And truthfully, it seemed to work because the teachers brought in for this reason felt like they had a voice in the process. None that I saw seemed to appreciate the depth of the whole process, so they all seemed to think they had made a difference in the development of the tests.
More specific to the author of the link, she seems like she's probably a good teacher, better than most. The education system, especially when it comes to statewide assessments, isn't prepared to appreciably handle outliers like her. She probably knows that. She probably knows the real battle to change these kinds of tests is not one she's prepared to tackle. I don't blame her. Nor do I blame her for making a public critique like she did.