> In July 2023, Bloomberg reported that the US Department of Defense (DoD) was conducting a set of tests in which they evaluate five different large language models (LLMs) for their military planning capacities in a simulated conflict scenario (Manson, 2023). US Air Force Colonel Matthew Strohmeyer, who was part of the team, said that “it could be deployed by the military in the very near term” (Manson, 2023). If employed, it could complement existing efforts, such as Project Maven, which stands as the most prominent AI instrument of the DoD, engineered to analyze imagery and videos from drones with the capability to autonomously identify potential targets. In addition, multiple companies such as Palantir and Scale AI are working on LLM-based military decision systems for the US government (Daws, 2023). With the increased exploration of the usage potential of LLMs for high-stakes decision-making contexts, we must robustly understand their behavior—and associated failure modes—to avoid consequential mistakes.