Foundations2026

Psychological Frameworks for Computational Modeling of Human Behavior

A Theoretical Foundation for Psychologically-Grounded Synthetic Populations

1. A Century of Discovery About Human Behavior

Every institution that has ever tried to organize people at scale has run into the same problem: predicting how they will behave. Armies needed to know which soldiers would hold their nerve under fire, hospitals which patients were likely to pull through, schools which students would actually learn.

Psychology took up that problem in the late nineteenth century and spent the following hundred years building a tested body of theory about how people think, feel, and act. Large language models now do something psychology never had to plan for: they simulate human responses at a scale no laboratory or survey could reach.

Whether that simulation runs on the same mechanisms this paper traces, or only reproduces their statistical surface in fluent language, is the real question here, and it can't be answered in general terms. It has to be answered layer by layer, against what a century of research actually found. That's where this section starts.

That century did not produce one answer. It produced eight, arriving in roughly the order that each new field of practice ran into a limit the last one couldn't explain.

Read separately, they can look like psychology changing its mind. Read together, they build on each other: each new finding narrowed what the next one had to explain, and none of the eight was ever overturned by the ones that came after. Section 2 turns that cumulative structure into an explicit architecture. First, though, come the discoveries themselves, in the order the century's practical problems forced them into view.

Discovery 1: Behavior follows laws

The industrial era needed to know whether behavior was lawful at all. Factories, militaries, and schools all had to train people at scale, and none of that works if behavior is wholly unpredictable. John Watson argued in 1913 that whatever happens inside the mind, what comes out is behavior, and behavior responds to its consequences in measurable ways. B. F. Skinner turned that claim into a research program. He showed that organisms repeat rewarded actions and abandon punished ones, and that unpredictable reward schedules produce the most persistent behavior of all — a finding from laboratory rats that later explained everything from gambling addiction to social media engagement. The framework worked. Skinner's adaptive teaching machines, Joseph Wolpe's systematic desensitization for phobias, and token economies that reorganized psychiatric wards all built on it. What it couldn't explain was just as concrete: identical conditioning still produces different behavior in different people. Behaviorism had no vocabulary for why.

Discovery 2: People differ systematically

Two world wars supplied the next question. Militaries needed to predict who would break under stress, hospitals who would recover, schools who would thrive. Answering that meant measuring what was inside people, not just how they behaved. Decades of factor-analytic research converged on five stable dimensions: Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism, now known as the Big Five. Each is heritable, measurable, and predictive of outcomes from career success to political orientation. A trait profile works like a statistical prior. Before any situation occurs, it already shapes the distribution of likely responses. Traits explained what behaviorism couldn't, but not everything. People with near-identical traits still diverge on the same situation, because a prior narrows a distribution without fixing an outcome.

Discovery 3: The mind has structure

Aviation, computing, and telecommunications made the next demand. Pilots needed instruments they could read under load, controllers needed to track several objects at once, and designers needed to know how much information a person could hold at all. The digital computer gave psychologists both a metaphor and a method. George Miller's 1956 paper found that working memory holds only a handful of items before earlier ones get displaced. Herbert Simon and Allen Newell built computational models of human problem-solving that could be tested against real performance data. Between them, the cognitive revolution delivered a set of hard constraints on information processing: capacity, decay, and the way prior knowledge shapes what gets perceived in the first place. What it missed was emotion, which it treated as noise degrading otherwise clean computation. A chess program and a human chess player were assumed to be running the same kind of process. They aren't.

Discovery 4: Situations can override character

The Second World War forced a harder question: how had millions of ordinary people taken part in systematic atrocity? Stanley Milgram's 1963 experiment tested it directly. Ordinary volunteers, believing they were assisting a study on learning, were told by an authority figure to deliver escalating shocks to a stranger who answered incorrectly. Sixty-five percent delivered what they believed were lethal shocks, simply on instruction, a result since replicated widely enough to stand on its own. Philip Zimbardo's Stanford prison study is often paired with Milgram's, but it carries less weight than its reputation suggests. Later scrutiny found that the guards' cruelty was substantially shaped by direct experimenter influence, and the study is now read as suggestive rather than confirmatory. The underlying claim doesn't need it. Milgram's result alone is enough to show that roles, norms, and hierarchy can override individual values in ways no cognitive model predicts. What it didn't explain was the emotional texture of the response: why the same situation feels threatening to one person and energizing to another.

Discovery 5: Emotion is mechanism, not noise

That question came out of the therapy room, not the laboratory. Clinicians treating depression, anxiety, and trauma kept finding that purely cognitive interventions failed: patients could see that a thought was distorted and still be unable to change it. Richard Lazarus's answer, developed through the 1960s, was that emotion isn't a reaction to an event but an evaluation of one. The same event produces fear in one person and excitement in another because they appraise it differently, against their own goals and their own sense of what they can handle.

Klaus Scherer, Ira Roseman, and the team of Ortony, Clore, and Collins spent the following decades specifying those dimensions precisely: not just whether an event is good or bad, but who caused it, whether it was expected, whether the person feels able to cope. Threat combined with low coping potential produces fear. Harm attributed to another person produces anger. A good outcome attributed to oneself produces pride. Given the appraisal, the emotion follows.

Antonio Damasio supplied the confirming case from the other direction. Patients who lost the capacity to feel emotion, through damage to the prefrontal cortex, kept their intellect fully intact and still made catastrophic everyday decisions. Emotion isn't noise interfering with judgment. It's part of how judgment works, supplying fast evaluative signals when deliberate calculation would be too slow.

This is the discovery that carries the most weight of the eight. It's the mechanism that turns a static trait profile into a specific response to a specific situation, and it becomes the hinge of the architecture in Section 2.

Discovery 6: Thinking operates in two modes

Economics had long assumed that people maximize expected utility under known probabilities. Daniel Kahneman and Amos Tversky spent the 1970s and 80s dismantling that assumption with data. People overweight rare events. They feel losses more sharply than equivalent gains. They anchor on numbers that carry no real information, and they judge frequency by how easily an example comes to mind. Their explanation, formalized as dual-process theory, holds that judgment runs on two systems: a fast, automatic one that produces most behavior, and a slow, effortful one that engages only under high stakes or real novelty, more often rationalizing what the fast system already decided than overriding it. Predicting behavior means knowing which mode is active, and that's a matter of arousal and cognitive load, not personality.

Discovery 7: Values structure goals

Cross-cultural research raised a different puzzle: why do people with similar personalities pursue different things across cultures? Shalom Schwartz found a universal structure of ten basic values — Self-direction, Stimulation, Hedonism, Achievement, Power, Security, Conformity, Tradition, Benevolence, and Universalism — arranged so that adjacent values reinforce each other and opposing ones pull apart. What differs across people and cultures is the weighting, not the structure. Values aren't goals. They're the criteria goals get measured against, defining what counts as a gain or a loss for a given person. Without this layer, appraisal has no way to answer its own central question — is this good or bad for me — because that question has no content until it's clear what the person is trying to achieve.

Discovery 8: Emotion is partly constructed

For decades, the dominant view held that emotions are fixed biological categories, expressed the same way everywhere. The evidence stopped supporting that. Brain imaging never turned up dedicated circuits for individual emotions, and different cultures categorize emotional experience differently. Lisa Feldman Barrett and the Theory of Constructed Emotion resolves the mismatch: emotion is actively built, in the moment, out of interoceptive signal, learned concepts, and cultural context (Barrett et al., 2007; Lindquist et al., 2012). The upshot is that emotional response is far more malleable, and far more dependent on context and prior belief, than a fixed-circuit account allows.

What a century of discovery produced

None of these eight discoveries was overturned by the ones that came after. Each added a layer the last one lacked the vocabulary to describe: stable traits, values, situational appraisal, cognitive constraints, social context, constructed emotion. What's missing, laid out this way, is an account of how they fit together into one system. That's what Section 2 builds.

2. A Layered Architecture

The eight discoveries above are usually taught as competing schools of thought. They read better as layers of one generative system, each constraining the layer below it.

Traits sit at the top: the stable individual differences captured by the Big Five. They don't produce behavior directly. They work as priors, shaping the distribution of likely responses before any situation occurs. Someone high in Neuroticism isn't fated to feel anxious in any given moment, but their baseline probability of threat appraisal is shifted.

Values sit beneath traits: the Schwartz structure of ten motivational priorities. Values convert a trait-shaped disposition into something appraisal can actually use, a criterion for what counts as a gain or a loss. Traits describe how a person tends to respond. Values describe what that person is protecting or pursuing. Neither layer alone can generate a specific response to a specific event.

That generation happens at the appraisal layer. Given a situation, a trait profile, and a value hierarchy, appraisal theory specifies the emotion that follows: threat with low coping potential produces fear, harm attributed to another produces anger, a goal-congruent outcome attributed to oneself produces pride. This is where upstream structure becomes a specific, first-person response. It's also the only layer that operates at the same grain as an actual decision — not a tendency, not a value, but this appraisal, in this moment, producing this action.

Two more layers modulate that response rather than sitting above or below it. Dual-process arbitration decides whether the fast, appraisal-driven response gets acted on directly or handed to slower deliberation, a function of arousal, cognitive load, and stakes rather than personality. Social context — conformity pressure, authority, an active group identity — can override individual appraisal entirely. That's why the same person, with the same traits and the same values, behaves differently in a crowd than alone.

Constructed emotion governs how the appraisal's output actually gets experienced and expressed, shaped by interoceptive state, learned concepts, and cultural framing. The same appraisal profile can surface as different emotion words, and different behavior, across individuals and cultures.

Put together, this isn't eight independent theories. It's one pipeline: trait priors, weighted into value-based goals, resolved by situational appraisal, modulated by arousal, load, and social context, and expressed as constructed emotional and behavioral output. A system that claims to simulate a person is implicitly claiming to instantiate all five stages. Most current approaches instantiate only the first.

LayerMechanismGovernsModeled as
TraitsBig Five factor structureBaseline response tendenciesStatic prior
ValuesSchwartz motivational circleWhat counts as gain vs. lossWeighted goal set
AppraisalLazarus / Scherer / OCC appraisal dimensionsSpecific emotion + action tendencyPer-situation function
ArbitrationDual-process (System 1 / System 2)Fast vs. deliberate responseState variable (arousal, load)
Social contextConformity, authority, group identityOverride of individual dispositionSituational modifier
Constructed emotionTheory of Constructed EmotionHow appraisal is experienced/expressedContext-dependent mapping

3. What This Requires of Computational Simulation

Section 2 makes a testable claim. A system that only encodes traits and values — a persona description, a set of demographic tags, a Big Five vector — has specified the prior, not the mechanism. It can describe a person's baseline tendencies. It can't generate their response to a specific, novel situation, because that response gets produced at the appraisal layer, and appraisal is a function of the situation, not of the person alone.

That yields a short list of concrete requirements.

A trait or demographic profile alone is not a simulation.

It's an input to one. Systems that stop here produce responses that look plausible in aggregate but stay static across situations: the same persona answering every question in the same register, regardless of what actually happened to it.

Appraisal has to be computed per situation, not retrieved from memory.

A system that has memorized how a Conscientious, Security-valuing person tends to respond will reproduce the statistical surface of that combination. A system that computes goal-relevance, attribution, and coping potential for this specific event is doing something closer to the real mechanism. This is testable: vary the situation while holding the persona fixed, and check whether the response shifts in the direction appraisal theory predicts, or only shifts stylistically.

Social context has to be able to override individual disposition.

A simulation with no representation of group identity, authority, or conformity pressure will systematically under-predict behavior in institutional and high-authority settings. Those happen to be the settings — workplaces, juries, classrooms, crowds — that most applied use cases care about most.

Arousal and cognitive load are state variables, not trait variables.

A person under time pressure or high stakes leans on fast, appraisal-driven response instead of deliberation. A simulation that always reasons deliberately is modeling the wrong system for most real decisions.

Emotional output is constructed, not retrieved.

The same underlying appraisal can surface as different labeled emotions depending on cultural and conceptual context. A system that hard-codes a fixed mapping from appraisal to emotion word will fail cross-culturally, and it will fail in predictable ways.

Together, these five requirements form a rough audit. Given any system that claims to simulate human behavior, ask which of the five pipeline stages it actually instantiates, and which it only approximates by pattern-matching on training data that already contains the statistical residue of these mechanisms without the mechanisms themselves. That distinction, mechanism versus residue, is answerable now in specific, checkable terms rather than as a rhetorical question.

4. Conclusion: Mechanism or Residue?

This paper opened by asking whether large language models actually model human psychology, or only reproduce its statistical surface in language. The pipeline in Section 2 and the audit in Section 3 turn that question from rhetorical into operational. A system can be tested layer by layer. Does its behavior shift with values as well as traits? Does its response to a new situation reflect goal-relevance and attribution, or only stylistic variation? Does it behave differently alone than in a group? Does it distinguish a snap judgment from a deliberated one? Does its emotional expression track cultural context, or run off a fixed lookup table?

For most systems built so far, the honest answer is mixed.

They're strong on the surface layers that show up richly in training data — trait language, value language, emotion words — and weak on the mechanism that connects them to a specific situation. That weakness isn't a reason to abandon the project. It's a specification for the next version of it. A synthetic population is only as psychologically grounded as the appraisal mechanism generating its responses, and unlike a persona description, that mechanism has to be built, tested, and validated against the century of research this paper has traced.

This is the challenge Heura Lab finds most compelling: building systems that don't just sound like people, but decide like them, layer by layer, the way this paper has laid out. It's early work, and there's a great deal left to test, but it's what we're excited to be building into our Psychologically Grounded Synthetic Population Platform and Decision Studio.

References

  • Barrett, L. F. (2017). How Emotions Are Made: The Secret Life of the Brain. Houghton Mifflin Harcourt.
  • Chomsky, N. (1959). Review of B. F. Skinner's Verbal Behavior. Language, 35(1), 26–58.
  • Costa, P. T., & McCrae, R. R. (1992). Revised NEO Personality Inventory (NEO PI-R) Professional Manual. Psychological Assessment Resources.
  • Damasio, A. R. (1994). Descartes' Error: Emotion, Reason, and the Human Brain. Putnam.
  • Haney, C., Banks, C., & Zimbardo, P. (1973). Interpersonal dynamics in a simulated prison. International Journal of Criminology and Penology, 1, 69–97.
  • Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
  • Lazarus, R. S., & Folkman, S. (1984). Stress, Appraisal, and Coping. Springer.
  • Lindquist, K. A., Wager, T. D., Kober, H., Bliss-Moreau, E., & Barrett, L. F. (2012). The brain basis of emotion: A meta-analytic review. Behavioral and Brain Sciences, 35(3), 121–143.
  • Milgram, S. (1963). Behavioral study of obedience. Journal of Abnormal and Social Psychology, 67(4), 371–378.
  • Miller, G. A. (1956). The magical number seven, plus or minus two. Psychological Review, 63(2), 81–97.
  • Ortony, A., Clore, G. L., & Collins, A. (1988). The Cognitive Structure of Emotions. Cambridge University Press.
  • Roberts, B. W., Kuncel, N. R., Shiner, R., Caspi, A., & Goldberg, L. R. (2007). The power of personality. Perspectives on Psychological Science, 2(4), 313–345.
  • Schwartz, S. H. (1992). Universals in the content and structure of values. Advances in Experimental Social Psychology, 25, 1–65.
  • Skinner, B. F. (1938). The Behavior of Organisms. Appleton-Century-Crofts.
  • Tajfel, H., & Turner, J. C. (1979). An integrative theory of intergroup conflict. In W. G. Austin & S. Worchel (Eds.), The Social Psychology of Intergroup Relations. Brooks/Cole.
  • Tversky, A., & Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185(4157), 1124–1131.
  • Wundt, W. (1897). Outlines of Psychology. Engelmann.