Educational systems, complexity science, and the evaluation of AI reasoning
Twenty-two years within one of the largest school districts in the United States, fourteen of them in leadership, a doctorate in urban educational leadership, an undergraduate degree in electrical engineering, and a research background in modeling how elementary classroom teacher-student interaction rules give rise to complex learning behavior. My work has consistently been the evaluation of judgment — in classrooms, in institutions, and now in machines.
My work in AI evaluation goes beyond identifying incorrect responses. I focus on finding incorrect structures: rules that execute out of sequence, instructions that quietly contradict one another, and contextual dependencies that others may overlook. Structural errors are the expensive kind. They survive testing, they repeat across every case the rule touches, and they are hardest to see once a system is already producing plausible output.
Fourteen years of school and district leadership, each position requiring consequential decisions under ambiguity and with incomplete information. That is precisely the discipline an evaluator exercises upon a model’s output.
Doctoral research treating classrooms as complex adaptive systems, together with a certificate in agent-based modeling from the University of Surrey. Agent-based modeling examines how elementary rules, executed in sequence, generate system-level behavior — a framework that transfers directly to the question of how a model’s instructions govern its conduct.
An undergraduate degree in electrical engineering, and the doctoral simulation written in NetLogo. As principal, I directed a school-wide one-to-one device initiative and secured district STEAM certification.
Adjunct professor within a university school-leadership program, following a decade of district professional development designed and delivered across three programs: instructional technology, data analysis to inform instruction, and content and developmental pedagogy. Composing instructions that hold in the order a learner encounters them is the entirety of that discipline.
Formal study in educational measurement and research methodology, chaos and complexity theory, organizational theory, and social network theory.
Learning assessment, curriculum design, structured feedback on written work, research synthesis, data interpretation, linear regression, and agent-based simulation programd in NetLogo for the doctoral research, together with daily working use of ChatGPT and Claude.
Native fluency in everyday and conversational Spanish, including register, tone and nuance a non-native reader would miss. My professional and technical vocabulary in Spanish is limited, so I do not take work requiring specialist terminology or translation.
| Qualification | Institution | Years |
|---|---|---|
| PhD, Educational Urban Leadership | Claremont Graduate University | 2011–2015 |
| Certificate, Agent-Based Modeling in the Social Sciences | University of Surrey, Guildford, United Kingdom | 2015 |
| MA, Educational Administration | California State University, Dominguez Hills | 2009–2010 |
| BS, Electrical Engineering (Psychology minor) | California State University, Long Beach | 1995–1999 |
| California Clear Teaching Credential | California State University, Los Angeles | 1999–2002 |
I examined classroom climate against mathematics achievement for the entire sixth grade of an urban middle school, 224 students, using linear regression. Of the seven dimensions of classroom climate measured, only two — Task-Orientation and Cohesiveness — proved statistically significant, and the climate variables together accounted for seven percent of the variance in achievement. The result is small, and it is reported as it stands.
I then constructed an agent-based simulation of those same classrooms in NetLogo, adapting two established formal accounts of how an individual learns and decides: the Rescorla-Wagner model of classical conditioning, and the Agent Zero model of neurocognitive decision-making. Both were modified to carry the seven climate dimensions, class size, classroom management structure and student learning thresholds. I calibrated the output against the observed achievement data, compared simulated behavior with observed behavior, documented the conditions under which the model failed, and specified the revisions warranting further testing.
That final step is the substance of the work. Producing a system that generates plausible output is the straightforward half. Establishing precisely where that output ceases to correspond to reality, and stating it without equivocation, is the half that carries weight.
Fields of interest: educational measurement and research methodology; chaos and complexity theory; organizational theory; global networks; social network theory; social learning theory; motivation and minority student achievement.
I write and publish fiction and essays through an imprint I own and operate. The relevance here lies not in the subject matter but in the discipline: the work is sustained, edited to a standard, and delivered on schedule — which is precisely what written evaluation demands.
Approximately 90,000 words each, written, edited and produced to fixed publication dates.
Short-form non-fiction, composed to a deadline and to a consistent voice.
Production, distribution, contracts, metadata and marketing across the full catalog.
Exhibited work, shown at ruthgamboagallery.com.