Behavioural Architectures and the Emergence of Strategic Games
A Conceptual Companion to A Mathematical Theory of Symbolic Selection: From Neurocognitive Interaction to Strategy-Space Compression and Loss-Dominant Equilibrium
This essay introduces the conceptual foundations of Symbolic Selection and its mathematical formalization. Its central proposition is that the games humans inhabit are not fixed, but emerge through repeated interaction between different behavioural architectures. Rather than assuming players, strategies, and payoffs already exist, it asks a more fundamental question: How do the games themselves come into being? Beginning with reciprocal and instrumental behavioural architectures, the essay explores how repeated interaction progressively shapes the relational environment in which future choices become possible. These ideas provide the conceptual foundation for The Psychopathic Selection Hypothesis (PSH), where instrumental behavioural architectures become increasingly adaptive within particular symbolic environments, and for the broader theory of Symbolic Selection, which examines how those symbolic environments recursively evolve over time. The accompanying mathematical paper formalizes these concepts, while the linked definitions provide more rigorous explanations for readers wishing to explore the framework in greater depth.
This essay focuses on the emergence of the relational game itself. The next essay examines what happens when these same dynamics scale beyond individuals. As repeated interactions accumulate across organizations, institutions, markets, and states, relational games become symbolic gameboards that increasingly shape the behaviour of everyone who enters them. That transition, from interpersonal dynamics to symbolic evolutionary systems, marks the beginning of Symbolic Selection as a general theory of institutional evolution and forms the bridge to the next stage of the mathematical framework.
Game theory has transformed our understanding of strategic behaviour. Across economics, evolutionary biology, political science, computer science, and artificial intelligence, it has provided an extraordinarily powerful mathematical language for analysing conflict, cooperation, negotiation, and competition. Despite its diversity, however, much of game-theoretic analysis begins from a common starting point. The game already exists. The players have already been identified, the rules governing interaction have already been established, the available strategies have already been defined, and the associated payoffs have already been specified. The central problem is therefore to determine how rational actors should behave within that strategic environment. This assumption is so deeply embedded within the discipline that it is rarely questioned. However, it raises what appears to be an even more fundamental problem. Before we ask how rational actors behave within a game, we might first ask how that particular game came into existence. Why does one environment repeatedly produce cooperation while another repeatedly produces manipulation? Why do some institutions stabilize around reciprocity while others stabilize around extraction? Before strategy can be analysed, the strategic environment itself requires explanation.
Consider the simplest possible example. Two individuals meet for the first time. Nothing of strategic significance yet exists between them. Neither owes the other anything. Neither possesses leverage. Neither has accumulated trust, obligation, resentment, dependency, reputation, or expectation. Every future remains possible because no history has yet constrained their interaction. They may cooperate, compete, exchange information honestly, deceive one another, negotiate, withdraw, or simply avoid further contact altogether. At this initial moment there is almost no meaningful game because there is almost no accumulated structure. Now imagine that these same individuals interact repeatedly over many years. Promises are either honoured or broken. Information is either shared or concealed. Favours accumulate. Debts accumulate. Trust either expands or deteriorates. Each interaction modifies the expectations governing the next. Gradually those expectations become sufficiently stable that both individuals begin behaving according to rules neither consciously designed. The strategic environment has changed because the history of interaction has changed. What did not exist when the relationship began now exists as an emergent property of accumulated interaction. The game has been constructed.
The Psychopathic Selection Hypothesis begins precisely at this point. Rather than asking how rational actors optimize behaviour within an existing game, it asks how repeated interaction gradually constructs different classes of games. This distinction is more than semantic. It shifts the primary object of analysis from strategy selection to game formation. The central claim is not that strategic environments should be treated as fixed analytical objects, but that they are themselves historical products generated through repeated interaction between heterogeneous behavioural architectures operating under particular symbolic boundary conditions. The question therefore changes from Which strategy is optimal? to What processes repeatedly produce the strategic environments within which optimization later occurs?
To approach this problem it becomes necessary to distinguish between behaviour and what PSH calls a behavioural architecture. Behaviour refers to individual actions. A behavioural architecture refers to the relatively stable computational organization responsible for generating those actions across time. It includes the manner in which an individual models other people, evaluates future outcomes, experiences emotional constraints, tolerates uncertainty, constructs objectives, and repeatedly selects among competing courses of action. Two individuals possessing equivalent intelligence and observing precisely the same external circumstances may nevertheless generate profoundly different behavioural trajectories because the internal organizations responsible for producing those behaviours differ. The important unit of analysis is therefore not the isolated action but the architecture capable of producing similar actions repeatedly across thousands of interactions.
This distinction immediately introduces another problem that conventional game-theoretic formulations generally leave implicit. Formal strategy spaces describe every strategy that is theoretically available within a game. Human beings, however, do not possess identical capacities to execute every formally available strategy. A reciprocal behavioural architecture may understand strategic deception perfectly while remaining unable to sustain it over extended periods because empathy, guilt, attachment, identity, or concern for long-term relationships continually interfere with execution. Conversely, an instrumental behavioural architecture may repeatedly engage in deception, strategic impression management, dependency formation, or coercive bargaining without encountering equivalent internal constraints. The distinction is therefore not simply which strategies exist objectively, but which strategies are psychologically realizable by different behavioural architectures. PSH refers to this distinction as the realizable strategy space. Every game contains an objective strategy space. Every behavioural architecture experiences only the subset it can reliably execute.
Whether these behavioural operations become adaptive depends not simply upon the actor but upon the environment within which repeated interaction occurs. PSH describes these environmental properties as symbolic boundary conditions because they define the adaptive landscape before the first interaction has taken place. These conditions include the scale of interaction, the persistence of reputation, the degree of anonymity, the speed with which consequences return to decision makers, the extent to which actions can be attributed to particular individuals, the ability to transfer costs onto others, the concentration of symbolic rewards such as money, status, or authority, the capacity to modify institutional rules, and the availability of meaningful alternatives for participants wishing to leave existing relationships. None of these variables determines behaviour directly. Rather, they determine which behavioural operations can accumulate successfully through repeated interaction and which encounter continual correction before becoming structurally significant. Whether a particular behavioural architecture succeeds therefore depends upon far more than the architecture itself. It depends upon the environment within which repeated interaction unfolds. A behavioural pattern that produces long-term success under one set of conditions may produce immediate failure under another. Evolutionary biology has understood this principle for generations. No biological trait is universally adaptive. Every trait exists in relation to an environment. Remove the environment and the concept of adaptation itself loses meaning. The same principle applies to symbolic systems. Behaviour cannot be evaluated independently of the conditions under which it is repeatedly expressed. The environment determines which behavioural operations accumulate, which are continually corrected, and which disappear before producing lasting structural consequences.
Consider the simple act of deception. Deception is not inherently adaptive or maladaptive. Its consequences depend almost entirely upon the surrounding symbolic environment. Within a small community characterized by repeated interaction, persistent reputation, immediate feedback, and limited anonymity, deception frequently destroys future opportunities. The immediate gain obtained through manipulation is often outweighed by the long-term loss of trust, cooperation, and social standing. The behavioural operation corrects itself because the environment continually returns the consequences of deception to the individual responsible for it. The same operation performed within a large symbolic institution may produce a very different result. Where interactions are anonymous, responsibility is distributed, consequences are delayed, and costs can be displaced onto individuals separated by geography, hierarchy, or time, deception may continue producing advantage long enough to become a stable feature of the surrounding system. The behaviour has not changed. The boundary conditions have.
This distinction leads naturally to a second concept that becomes central to PSH: the difference between an objective strategy space and a realizable strategy space. Classical game-theoretic models typically begin by defining every strategy available within a particular game. From a mathematical perspective, every participant appears capable of selecting any of those strategies. Human behaviour rarely operates in this manner. The formal existence of a strategy does not imply that every behavioural architecture can repeatedly execute it. A reciprocal behavioural architecture may fully understand manipulation while remaining unable to sustain it because empathy, guilt, attachment, or moral identity continually interfere with execution. An instrumental behavioural architecture may repeatedly execute precisely the same behaviour without encountering those constraints. The game therefore contains an objective strategy space, while each behavioural architecture experiences only the subset of strategies it can repeatedly realize. Adaptation depends not merely upon what strategies exist, but upon the interaction between the realizable strategy space of the actor and the symbolic boundary conditions of the environment.
This distinction begins to explain why repeated interaction does not simply produce different outcomes; it produces different histories. Every interaction alters the conditions under which subsequent interactions occur. A promise that is honoured increases the probability that future promises will be believed. A deception that remains undetected increases the probability that deception will be attempted again. A favour creates obligation. An unpaid debt creates dependency. A fair exchange increases confidence in future cooperation. None of these changes remains isolated. They accumulate. The strategic relationship existing after ten thousand interactions differs fundamentally from the relationship that existed after the first because each interaction leaves behind constraints, expectations, opportunities, and obligations that did not previously exist.
At this point the concept of a game begins to change. Rather than treating the game as a predefined object, PSH treats it as the accumulated structure produced through repeated interaction. A game is not simply a collection of formal rules. It is the historically constructed network of expectations, dependencies, incentives, constraints, and opportunities inherited by participants before they make their next decision. Every previous interaction contributes to that inherited structure. Some interactions expand the autonomy of future participants by increasing trust, preserving meaningful alternatives, and encouraging reciprocal exchange. Other interactions progressively reduce autonomy by concentrating information, increasing dependency, eliminating alternatives, and making future behaviour increasingly predictable. The game is therefore not static. It is continuously rewritten by the behavioural operations repeatedly executed within it.
The distinction between reciprocal and instrumental behavioural architectures now acquires greater significance. These terms should not be understood as moral categories, nor as descriptions of good and evil. They describe different computational approaches to uncertainty. Reciprocal architectures primarily reduce uncertainty by stabilizing relationships. Trust, transparency, reciprocity, conflict repair, and the preservation of mutual autonomy reduce the need for continual strategic calculation because future behaviour becomes increasingly predictable through cooperation itself. Instrumental architectures solve the same problem differently. Rather than stabilizing relationships through mutual trust, they reduce uncertainty by increasing asymmetries of knowledge, creating dependency, narrowing alternatives, restructuring incentives, and improving their ability to anticipate or constrain the behaviour of other actors. Both architectures attempt to reduce uncertainty. They simply employ different behavioural operations to accomplish that objective.
The difference between these behavioural architectures does not lie simply in the behaviours they perform, but in the behavioural protocols they repeatedly execute throughout the life of a relationship. A behavioural protocol (as defined in the PSH framework) should not be understood as a single behaviour. It is an ordered sequence of behavioural operations deployed across successive interactions to achieve a particular objective. A reciprocal behavioural protocol may include operations such as honest signalling, reciprocal exchange, emotional attunement, fulfilled commitments, transparent communication, conflict repair, and the preservation of meaningful alternatives for the other participant. An instrumental behavioural protocol may instead include operations such as social modelling, mimicry, impression management, selective reinforcement, strategic deception, dominance seeking, alternative suppression, information control, and dependency formation. The effectiveness of either protocol depends not only upon the operations themselves but upon their timing, their order, and the changing cognitive and emotional state of the agents participating in the relationship. The same behavioural operation performed during the first interaction may produce no meaningful consequence, while the identical operation performed after hundreds of previous interactions may fundamentally alter the trajectory of the relationship. Behavioural protocols are therefore inherently recursive. Every operation modifies the conditions under which subsequent operations will later be interpreted.
To understand why this occurs, it is necessary to recognize that human beings do not respond directly to objective reality. Every agent continually constructs an internal cognitive and affective model through which reality is interpreted. This internal model contains expectations concerning the future behaviour of other people, beliefs about oneself, emotional attachment, trust, perceived safety, anticipated consequences, perceived alternatives, and estimates of what actions remain realistically possible. Decisions emerge from this internal model rather than directly from the external environment. Every interaction therefore becomes an opportunity for behavioural operations to modify not merely behaviour, but the internal model responsible for generating future behaviour.
Reciprocal behavioural protocols exploit this process in a manner that is generally adaptive within cooperative environments. Repeated honest communication, fulfilled commitments, emotional attunement, reciprocal exchange, and conflict repair gradually increase the correspondence between each participant’s internal model and the behaviour actually exhibited by the other. Prediction becomes progressively more accurate because repeated interaction continually confirms rather than violates previous expectations. Emotional attachment deepens because the relationship repeatedly demonstrates reliability rather than contradiction. Trust therefore emerges not simply because people feel positively toward one another, but because their internal models become increasingly capable of predicting one another’s future behaviour with comparatively little uncertainty. The relationship becomes computationally efficient because less continual vigilance is required to anticipate future interaction. Reciprocal attachment is therefore not a weakness. It is an adaptive consequence of repeatedly interacting within an environment where cooperation is both genuine and mutually beneficial.
Instrumental behavioural protocols follow a fundamentally different trajectory. Their objective is not merely to influence immediate behaviour but to progressively reorganize the internal model maintained by the other participant. This reorganization cannot occur through coercion alone. At the beginning of a relationship there exists neither sufficient attachment nor sufficient epistemic authority for coercive operations to produce lasting change. Consequently, instrumental protocols frequently begin with operations that establish precisely those conditions upon which later influence depends. Social modelling creates perceived similarity. Mimicry produces familiarity. Impression management increases perceived competence, desirability, or trustworthiness. Selective reinforcement strengthens emotional attachment by repeatedly associating the relationship with safety, validation, excitement, or relief from uncertainty. Each operation leaves the external relationship apparently intact while simultaneously modifying the internal model through which the reciprocal agent interprets that relationship.
Note: It is important to emphasize that the instrumental behavioural architecture described here is grounded in decades of psychopathy research rather than moral characterization. Across the work of Robert Hare, Kent Kiehl, Christopher Patrick, James Blair, and others, psychopathy is consistently associated with a strategic orientation toward reward acquisition, dominance, control, and the instrumental use of other people. Individuals high in these traits often retain intact cognitive empathy (the ability to understand and model the thoughts and emotions of others) while exhibiting markedly reduced affective empathy, guilt, or remorse when causing harm. Deception, manipulation, impression management, and strategic exploitation therefore become psychologically realizable strategies because they encounter comparatively few internal emotional constraints. Reciprocal behavioural architectures, by contrast, remain constrained by empathy, attachment, guilt, concern for others, and the preservation of mutually beneficial relationships. The distinction developed throughout this essay is therefore not one of intelligence or competence, but of behavioural architecture: one seeks mutual adaptation through reciprocity, while the other seeks strategic advantage through the instrumental organization of relationships. This distinction forms one of the central conceptual foundations of The Psychopathic Selection Hypothesis (PSH).
Only after this reorganization has occurred do later operations acquire their full effectiveness. Criticism expressed by a stranger possesses relatively little influence because the stranger occupies no privileged position within the recipient’s internal model. The same criticism expressed by an individual who has gradually become incorporated into that model may profoundly alter self-evaluation, emotional stability, and future behaviour. Statements such as “You’ll never find anyone else,” “You’re worthless,” or “Everything wrong in this relationship is your fault” do not function merely as information. They become increasingly influential because the target has learned, through hundreds of previous interactions, to assign emotional significance and epistemic credibility to the individual delivering them. The protocol therefore operates recursively. Earlier behavioural operations reorganize the target’s internal model in ways that allow later operations to become progressively more effective.
Dependency emerges through this recursive process rather than through any single act of domination. The reciprocal agent does not become dependent because autonomy suddenly disappears. Rather, successive behavioural operations gradually reorganize the internal model through which future decisions are made. Emotional attachment deepens. Identity becomes increasingly intertwined with the relationship. Confidence in independent judgment weakens as increasing epistemic authority is assigned to the other participant. Self-evaluation becomes progressively mediated through external validation originating from within the relationship. Perceived alternatives become less psychologically available, not necessarily because they have objectively vanished, but because the internal model now evaluates them as increasingly improbable, undesirable, or threatening. Behaviour consequently reorganizes itself around preserving the relationship because the relationship has become integrated into the very processes through which the agent predicts the future, regulates emotion, evaluates self-worth, and pursues personally meaningful objectives.
The relationship has therefore undergone a qualitative transformation. At its beginning, departure represented little more than the termination of a social interaction. After hundreds or thousands of recursive updates to the internal models of both participants, departure requires abandoning a cognitive and emotional framework through which identity, belonging, security, expectation, and future possibility have gradually become organized. Dependency is therefore neither purely emotional nor purely rational. It is an emergent property of neurocognitive co-evolution. Two behavioural architectures have repeatedly acted upon one another, yet they have not done so symmetrically. One architecture has continually updated its internal model in good faith, assuming reciprocal cooperation. The other has repeatedly deployed behavioural operations designed to reorganize that very model. The result is not merely a different relationship. It is a different game.
Note: The sequence described here should not be understood as ending with dependency formation. Within both interpersonal relationships and larger symbolic systems, dependency often functions as the preferred strategy precisely because it minimizes resistance while maximizing compliance. However, when pacification fails, the available behavioural repertoire frequently expands to include more overt forms of coercion. Reputational attack, intimidation, threats, exclusion, legal pressure, economic punishment, surveillance, and ultimately physical violence can all be understood as later-stage behavioural protocols deployed when earlier strategies no longer secure compliance. Violence is therefore not the defining characteristic of instrumental behavioural architectures; it is often the least efficient one. Consistent with the broader argument developed throughout PSH, stable systems of extraction generally prefer voluntary participation, dependency, and perceived legitimacy over overt coercion because these approaches are less costly, more scalable, and more durable. Punishment frequently appears only after reciprocal compliance has failed.
An instrumental behavioural architecture does not ordinarily begin a relationship through coercion, domination, or overt exploitation. Such behavioural operations would almost certainly fail because the reciprocal agent has not yet incorporated the instrumental actor into their internal model of reality. At the beginning of the interaction, the instrumental actor possesses neither emotional significance nor epistemic authority. Their opinions carry little weight, their criticisms are easily dismissed, and their demands can simply be refused (as stated). Consequently, instrumental behavioural protocols begin elsewhere. Their immediate objective is not extraction. Their immediate objective is to establish the conditions under which future influence becomes possible.
These protocols typically begin with behavioural operations such as social modelling, mimicry, impression management, strategic self-disclosure, selective reinforcement, and strategic deceit. Although these operations differ in their immediate execution, they perform a common computational function. They progressively reshape the reciprocal agent’s internal representation of the instrumental actor, increasing perceived similarity, trustworthiness, competence, desirability, credibility, and emotional significance. The protocol therefore operates not by immediately changing the reciprocal agent’s behaviour, but by changing the model through which future behaviour will later be generated.
Social modelling illustrates this process particularly well. The instrumental actor first gathers information regarding the reciprocal agent’s preferences, values, ambitions, fears, insecurities, emotional needs, political beliefs, humour, interpersonal style, and prior relational experiences. This information is not collected passively. It provides the behavioural material from which subsequent protocol execution is constructed. The instrumental actor then selectively reflects these characteristics back to the reciprocal agent. A favourite novel unexpectedly becomes their favourite novel. Musical preferences appear identical. Political values converge. Life goals seem remarkably compatible. Personal experiences appear strangely familiar. Whether these representations correspond to the instrumental actor’s genuine preferences is secondary. Their function is to increase perceived similarity because reciprocal behavioural architectures naturally interpret similarity as evidence of compatibility, shared identity, and reduced interpersonal uncertainty.
Impression management performs a related but distinct operation. Rather than increasing perceived similarity, it selectively modifies the reciprocal agent’s evaluation of the instrumental actor’s status, competence, desirability, credibility, or moral character. Educational achievements may be exaggerated. Professional accomplishments embellished. Financial stability overstated. Emotional maturity carefully performed. Acts of generosity selectively displayed while failures, inconsistencies, or contradictory behaviours remain concealed. Strategic deceit frequently operates within this same process, not necessarily through elaborate fabrication but through selective disclosure, omission, exaggeration, or the careful presentation of information that encourages the reciprocal agent to construct conclusions advantageous to the instrumental actor. The objective is not simply to communicate false information. It is to recursively shape the model through which the reciprocal agent understands who this individual is.
The reciprocal behavioural architecture responds to these operations in precisely the manner it evolved to respond within genuinely cooperative environments. Repeated experiences of perceived similarity strengthen interpersonal trust. Apparent consistency between words and behaviour increases credibility. Positive emotional experiences deepen attachment. Successful prediction of the instrumental actor’s behaviour increases confidence that the relationship is both safe and mutually beneficial. Gradually, the reciprocal agent begins assigning increasing epistemic authority to the instrumental actor. Their observations become increasingly informative. Their advice acquires greater influence. Their evaluations carry greater emotional weight. This process should not be regarded as a cognitive failure. It is the ordinary neurocognitive mechanism through which intimacy, friendship, family formation, scientific collaboration, and cooperative social life become possible. Human beings do not merely exchange information within close relationships. They gradually participate in the joint construction of reality itself.
This transition fundamentally alters the relationship. At its beginning, the instrumental actor exists as one opinion among many. Their evaluations possess no privileged status within the reciprocal agent’s internal model. As attachment deepens, however, the relationship itself increasingly participates in the reciprocal agent’s processes of identity formation, emotional regulation, self-evaluation, and future planning. The instrumental actor is no longer experienced simply as another person. They become one of the reference points through which the reciprocal agent interprets themselves and the surrounding world.
Only after this neurocognitive reorganization has occurred do later behavioural operations acquire their full effectiveness. Criticism expressed by a stranger can usually be dismissed because the stranger occupies no meaningful position within the reciprocal agent’s construction of reality. The same criticism expressed by a deeply trusted partner produces an entirely different response. The reciprocal agent does not merely hear the statement. They begin asking whether it reveals something fundamentally true about themselves. Behavioural operations that initially possessed little influence now acquire extraordinary power because they are interpreted through an internal model that has been recursively reorganized by hundreds of previous interactions. The protocol therefore becomes self-reinforcing. Earlier behavioural operations create the conditions under which later behavioural operations become progressively more effective.
Only now does dependency begin to emerge. It does not arise because the reciprocal agent suddenly loses autonomy, nor because external alternatives have physically disappeared. It emerges because the reciprocal agent increasingly relies upon the relationship as a source of emotional regulation, self-understanding, future orientation, and epistemic guidance. Independent judgment is exercised less frequently, not because independent thought has become impossible, but because another behavioural architecture has gradually become incorporated into the reciprocal agent’s own processes of evaluating reality. What began as adaptive interpersonal trust has become progressively transformed into epistemic dependence. The reciprocal agent increasingly experiences both themselves and the future through a model of reality that has been recursively shaped by another agent pursuing a different objective function.
The behavioural protocols developed in the previous section should not be understood as isolated techniques of interpersonal influence. They are successive operations directed toward a common optimization problem. Having progressively reorganized the reciprocal agent’s internal model through trust, attachment, epistemic authority, and dependency formation, the instrumental behavioural architecture confronts a different challenge. The reciprocal agent remains autonomous. They retain the capacity to reinterpret the relationship, seek independent sources of information, revise previously held beliefs, refuse cooperation, expose deception, or leave altogether. Continued access to the instrumental actor’s desired outcomes therefore remains fundamentally uncertain.
This uncertainty constitutes what PSH describes as the Target Autonomy Problem. The problem arises whenever one agent’s objective function depends upon the future behaviour of another autonomous agent. Whether the desired outcome is affection, admiration, sexual access, labour, money, information, status, emotional regulation, compliance, political support, or some other valued outcome is ultimately secondary. The underlying computational problem remains identical. Future reward depends upon another human being who retains the capacity to generate behaviour inconsistent with the instrumental actor’s objectives. The problem is therefore not simply how to influence behaviour once. It is how to progressively reduce uncertainty concerning future behaviour while the target continues experiencing themselves as acting voluntarily.
Importantly, the Target Autonomy Problem is not solved by eliminating autonomy. Human beings cannot ordinarily be controlled with mechanical precision, nor is such precision required. Rather, the problem is solved by progressively reorganizing the cognitive, emotional, and relational conditions under which autonomy is exercised. The reciprocal agent continues making decisions throughout the relationship. What changes are the neurocognitive processes through which those decisions are generated. Every recursive update to the reciprocal agent’s internal model subtly alters how future situations are interpreted, how emotional significance is assigned, how consequences are anticipated, how competing alternatives are evaluated, and ultimately which behavioural strategies appear psychologically realizable. The instrumental actor therefore does not primarily shape individual decisions. The protocol progressively reshapes the decision landscape itself.
This account is broadly consistent with several well-established findings across psychology and cognitive neuroscience. Attachment research demonstrates that close relationships become integrated into systems of emotional regulation and self-concept. Research on epistemic trust shows that human beings naturally organize social learning around individuals perceived as reliable, familiar, or trustworthy, allowing those individuals to acquire increasing influence over how information is interpreted. Studies of coercive control similarly demonstrate that repeated degradation, conditional reinforcement, isolation, and manipulation of self-worth can progressively alter self-evaluation, perceived alternatives, and decisions regarding resistance or departure. None of these literatures, however, provides a formal game-theoretic account of how these processes recursively alter strategic interaction through time. The contribution of PSH is not the observation that attachment influences behaviour, that trust shapes decision-making, or that coercive relationships produce dependency. These phenomena are already well documented. The contribution of PSH is to organize these findings into a recursive computational framework in which ordered behavioural protocols progressively modify internal cognitive models, relational states, realizable strategy spaces, and ultimately the games from which future behaviour emerges.
This distinction also clarifies where the present framework departs from conventional game-theoretic analysis. Classical game theory has been extraordinarily successful in analysing strategic interaction under specified rules, payoff structures, and strategy spaces. Within that framework, agents optimize over a game that is assumed to exist prior to analysis. The origins of behavioural architectures, the neurocognitive processes through which strategies are generated, and the historical mechanisms through which games themselves evolve ordinarily remain outside the formal model. These are not shortcomings of classical game theory. They reflect its intended level of abstraction. PSH begins precisely where those assumptions end. Rather than assuming fixed games, stable strategy spaces, and static objective functions, it asks how repeated behavioural protocols recursively modify the internal cognitive models of interacting agents, how those modifications alter the relational variables from which future strategic interaction is generated, and how those accumulated changes progressively reconstruct the game itself. The present framework therefore does not propose an alternative solution concept for an existing game. It proposes a mechanistic account of how games themselves are historically constructed before conventional optimization begins.
At this point, a distinction introduced earlier becomes indispensable. The objective strategy space consists of every action formally available within the game. A reciprocal agent may remain objectively capable of resisting, negotiating, seeking external advice, exposing deception, terminating the relationship, or pursuing an entirely different future. These possibilities continue to exist within the formal structure of the interaction. The realizable strategy space, however, consists only of those strategies the agent can genuinely perceive, emotionally tolerate, evaluate as achievable, and successfully execute given their current behavioural architecture, internal model, relational state, and surrounding environment. Human beings do not optimize over every objectively available possibility. They optimize over the subset of possibilities their evolving neurocognitive model represents as psychologically and behaviourally realizable.
The significance of the behavioural protocols now becomes fully apparent. Their cumulative objective is not necessarily to eliminate alternatives from the external world. Rather, it is to recursively reorganize the reciprocal agent’s evaluation of those alternatives. A future without the relationship may remain objectively possible while becoming emotionally unimaginable. Independent judgment may remain intellectually available while feeling increasingly unreliable. Friends, employment, or alternative relationships may continue to exist while becoming progressively associated with uncertainty, anticipated rejection, shame, instability, or failure. The objective structure of the game may therefore remain largely unchanged even as the reciprocal agent’s lived experience of that game undergoes profound transformation. The behavioural architecture has not lost objective freedom. It has progressively lost access to the strategies through which that freedom could previously be realized.
It is this recursive contraction of the realizable strategy space that gives rise to the next stage of the theory. Strategy-space compression does not begin when alternatives disappear from the external world. It begins when repeated interaction progressively reorganizes the reciprocal agent’s internal model such that an increasing number of objectively available futures cease to exist as psychologically realizable possibilities. Formal freedom and practical freedom begin to diverge. The reciprocal agent remains autonomous in principle while experiencing progressively fewer futures as genuinely livable. It is from this divergence that entrapment, and ultimately Loss-Dominant Equilibrium, begin to emerge.
Every successful execution of an instrumental behavioural protocol changes the conditions under which the next interaction begins. The reciprocal agent does not return to a neutral state after each exchange. Attachment, trust, self-evaluation, anticipated loss, perceived alternatives, and expectations regarding future behaviour are carried forward into the next decision. The relationship therefore acquires memory. No interaction occurs within precisely the same game as the interaction that preceded it because the internal and relational variables governing choice have already changed. The game inherits its structure from its own history.
Dependency is one expression of this accumulated history, but it is not a single psychological state. It is a relational condition produced when increasingly important functions become organized through continued access to the relationship. Emotional regulation may become dependent upon the instrumental actor’s approval or presence. Self-evaluation may become increasingly sensitive to their judgments. Identity may become organized around the shared relationship and its anticipated future. Social belonging, housing, income, reputation, sexual intimacy, professional access, or family continuity may also become tied to continued participation. These forms of reliance can reinforce one another. The emotional cost of departure rises because material security is threatened; the material cost feels still greater because confidence in independent action has weakened; the loss of identity appears catastrophic because alternative relationships and futures have become increasingly difficult to imagine.
This accumulation changes the way strategies are evaluated. A strategy is not selected merely because it exists. It must first appear sufficiently intelligible, tolerable, and achievable to enter serious consideration. Human decision-making is shaped by anticipated emotion, perceived self-efficacy, attachment, threat, social belonging, expected regret, and sensitivity to loss. An action associated with overwhelming shame, fear, grief, uncertainty, or anticipated failure may remain objectively available while no longer functioning as a practical option. The reciprocal agent may understand, at an abstract level, that leaving is possible while being unable to construct a believable model of life after departure. The difficulty is not simply choosing between two known outcomes. It is generating a psychologically coherent future in which the alternative can be imagined as survivable.
The reciprocal agent’s internal model therefore affects strategy generation before explicit comparison begins. An individual who has repeatedly been told that they are incompetent may no longer generate independent employment as a credible possibility. An individual whose friendships have been systematically devalued may no longer represent those friendships as reliable sources of support. An individual whose self-worth has become tied to the instrumental actor’s approval may experience withdrawal not merely as rejection but as evidence confirming personal worthlessness. The objective possibilities have not necessarily vanished. The generative model from which possible futures arise has changed.
This is the neurocognitive dimension of strategy-space compression. The brain does not evaluate an exhaustive catalogue of every action theoretically available in the world. It constructs a limited set of salient possibilities from prior experience, learned expectations, emotional significance, perceived capability, and current threat. Repeated behavioural protocols can modify each of these inputs. Strategic criticism alters perceived capability. Information control alters what evidence is available. Alternative suppression reduces exposure to competing models of the relationship. Conditional reinforcement changes the anticipated emotional consequences of compliance and refusal. Commitment escalation attaches prior investments to continued participation. As these updates accumulate, strategies that once emerged readily during deliberation may cease to be generated, while others may appear so costly that they are rejected before meaningful evaluation occurs.
The game-theoretic importance of this process lies in the distinction between formal availability and effective availability. Classical models can represent exit as one element of the player’s strategy set. Yet the formal presence of exit does not establish that the actor can realize it under the conditions produced by the game’s history. PSH therefore treats the realizable strategy space as a dynamic object. It changes with the actor’s behavioural architecture, internal model, objective function, relational state, and surrounding constraints. A strategy can remain inside the objective strategy space while moving outside the realizable strategy space because its perceived cost has risen, its anticipated probability of success has fallen, or the actor can no longer generate a coherent trajectory through which it might be executed.
Strategy-space compression may occur through external and internal pathways at the same time. External compression occurs when the relationship materially removes alternatives. Access to money may be restricted. Social contacts may be severed. Housing may depend upon continued participation. Threats may raise the cost of disclosure. Shared children, legal obligations, debt, or professional dependence may create genuine constraints. Internal compression occurs when the reciprocal agent’s evolving model assigns greater danger, lower expected success, or less psychological tolerability to the alternatives that remain. The two pathways interact. A small material constraint may exert far greater influence once self-confidence has deteriorated, while a belief in personal incapacity becomes more credible when real resources and support have also been reduced.
The process need not involve the complete disappearance of any strategy. Compression can occur when exit, resistance, refusal, retaliation, or support-seeking remain possible but become increasingly costly relative to compliance. The agent’s strategy space may therefore contract in size, or it may retain the same formal elements while becoming sharply distorted by changing cost gradients. A strategy that was once ordinary becomes exceptional. A strategy that was once uncomfortable becomes terrifying. A strategy that was once easily reversible becomes associated with irreversible loss. The practical structure of choice changes even where the legal or physical structure appears unchanged.
Consider resistance. Early in the relationship, disagreement may carry little consequence. The reciprocal agent can refuse a request, challenge an interpretation, or step away without threatening the foundations of identity or security. After dependency has deepened, the same act of resistance may threaten emotional withdrawal, humiliation, retaliation, financial instability, or the loss of a shared future. Resistance is still possible, but it now carries a different anticipated cost. The strategy has not been removed from the game. Its location within the decision landscape has changed.
The same is true of seeking support. At the beginning, the reciprocal agent may discuss the relationship openly with friends or family. Later, disclosure may require admitting that earlier judgments were mistaken, confronting shame, risking disbelief, or contradicting the instrumental actor’s repeated portrayal of outside relationships as hostile or untrustworthy. The agent may anticipate being blamed for remaining, pressured to leave before feeling ready, or forced to defend a relationship whose earlier form still remains emotionally present. Seeking support therefore becomes more than a simple informational act. It threatens the shared reality through which the relationship has been maintained.
Exit undergoes the most extensive transformation. At the outset, departure means ending a developing relationship. Later, it may mean losing a partner, home, social role, financial arrangement, imagined future, sense of competence, source of emotional regulation, and one of the principal reference points through which the agent understands themselves. It may also require accepting that the relationship presented at the beginning never existed in the form in which it was believed. The anticipated loss is therefore not limited to the present relationship. It includes the collapse of the past narrative that justified prior investment and the future narrative around which current identity has been organized.
Commitment escalation deepens this problem because previous investments alter the meaning of reversal. Time, affection, money, sacrifice, reputation, promises, and identity have already been committed. Leaving can therefore be experienced not merely as accepting future uncertainty but as declaring those investments unrecoverable. The agent may continue not because further investment is likely to restore the relationship, but because departure would make the accumulated loss immediate and undeniable. Hope serves an important function here. The reciprocal agent may continue orienting toward the earlier version of the relationship, interpreting periods of warmth or remorse as evidence that the original bond can be recovered. Intermittent reinforcement preserves this possibility by preventing the internal model from converging completely on the conclusion that the relationship is irreversibly harmful.
Strategy-space compression is therefore not simply imposed upon a passive target. It emerges through the interaction between the instrumental actor’s protocols and the reciprocal actor’s adaptive capacities. Attachment, hope, forgiveness, openness to revision, sensitivity to relationship loss, and commitment to repair are not pathological traits. They are central to durable human cooperation. Under reciprocal conditions, they allow relationships to survive conflict, error, illness, and temporary instability. Under instrumental conditions, the same capacities can sustain continued participation while the relational structure becomes progressively less reciprocal. The target remains partly because the architecture that makes genuine commitment possible also resists abandoning a relationship at the first sign of failure.
At a certain point, however, dependency and strategy-space compression produce a further qualitative change. The reciprocal agent may recognize that the relationship is harmful. They may accurately identify deception, control, degradation, or the continuing loss of autonomy. Recognition alone does not restore the strategies that have become difficult to realize. Knowledge that one should leave is not equivalent to possessing the emotional, cognitive, social, and material capacity to do so. The actor can therefore understand the long-term direction of the relationship while remaining unable to act upon that understanding within the immediate decision horizon.
This is the condition PSH describes as entrapment. Entrapment occurs when exit becomes locally more costly than continued participation even though continued participation produces greater harm over the longer term. The reciprocal agent is not necessarily confused about the quality of the relationship. The conflict lies between time horizons. Staying distributes harm gradually and preserves elements of attachment, stability, identity, or security in the present. Leaving concentrates loss immediately. It may produce grief, fear, financial disruption, social exposure, uncertainty, retaliation, or the collapse of the future around which the agent has organized their life.
The distinction between local and long-term evaluation is essential. From an external perspective, exit may clearly offer greater future welfare. The observer sees the relationship’s cumulative harm and compares it against the possibility of recovery. The reciprocal agent must instead cross the immediate loss barrier produced by the existing relational state. They must act before the benefits of exit are available and while its costs are most concentrated. The better long-term strategy may therefore remain locally dominated by the strategy of continuation.
Entrapment should not be mistaken for the complete absence of agency. The reciprocal agent continues to make decisions, bargain, resist, accommodate, seek temporary safety, and preserve parts of the self. Yet these actions occur within a field whose most consequential alternatives have become difficult to realize. The agent may optimize carefully inside that field while remaining unable to transform the field itself. What appears from outside as passivity may consist internally of continuous adaptation to a narrowing set of survivable possibilities.
The game has now acquired a stable structure that neither participant faced at the beginning. The instrumental actor has obtained more predictable access to valued outcomes because resistance and exit carry greater costs. The reciprocal actor continues participating because the available deviations concentrate anticipated loss. The relationship can therefore persist without producing mutual benefit and without either participant believing that it represents the best possible future.
This is the basis of what PSH calls a Loss-Dominant Equilibrium. In a conventional equilibrium, stability is often explained by the absence of a unilateral deviation that would improve the player’s payoff. In a Loss-Dominant Equilibrium, the arrangement remains stable for a different reason. Deviation may improve long-term welfare, but it produces a larger and more concentrated loss in the present. Continuation is selected not because it is desirable, but because the realizable alternatives appear immediately worse.
The reciprocal agent is therefore still optimizing, but the structure of optimization has changed. At the beginning of the relationship, decisions were organized around the pursuit of positive possibilities: intimacy, cooperation, belonging, security, and a shared future. Under entrapment, decisions become organized around avoiding the most threatening loss. Compliance may preserve attachment. Silence may prevent escalation. Staying may preserve housing, identity, access to children, or the hope of relational repair. None of these outcomes represents genuine flourishing. They are strategies for preventing an anticipated collapse that appears even more dangerous than continued harm.
Loss dominance explains how a damaging game can remain stable without being beneficial. The equilibrium is not maintained by shared welfare, equal power, or successful cooperation. It is maintained because the history of the game has attached concentrated loss to deviation. The reciprocal agent may understand that continuation is destructive while still experiencing departure as intolerable. The instrumental actor may obtain short-term reward while progressively degrading the relationship’s long-term viability. Local strategic success and total relational failure can therefore coexist.
At this point the minimal relational game is complete. Ordered behavioural protocols have modified the reciprocal agent’s internal cognitive and affective model. Those updates have altered trust, attachment, dependency, anticipated loss, exit cost, and perceived alternatives. These relational changes have compressed the reciprocal agent’s realizable strategy space. Compression has produced entrapment by making departure locally irrational despite long-term harm. Entrapment has stabilized a Loss-Dominant Equilibrium in which continued participation minimizes immediate loss rather than maximizing welfare.
The game present at the end of this process is not the game that existed at the beginning. It contains different expectations, different constraints, different effective strategies, and a different distribution of predictability and control. No external designer needed to specify its final form. It was constructed recursively through the accumulated history of interaction.
The remaining question is whether this functional sequence remains confined to relationships between individuals. If protocols can modify internal models, relational states, realizable strategy spaces, and payoffs within the minimal game, under what conditions might functionally equivalent operators appear within organizations and institutions? That question belongs to the problem of symbolic scaling. It begins only after the minimal relational mechanism has been established.
The argument developed throughout this essay has remained deliberately confined to the minimal relational game and how such games are emergent from human behaviour itself. Beginning from the distinction between reciprocal and instrumental behavioural architectures, we proposed that repeated behavioural protocols recursively modify the internal cognitive and affective models of interacting agents. These recursive updates reorganize trust, attachment, epistemic authority, dependency, and anticipated loss, progressively altering the relational variables from which future strategic interaction is generated. As these relational variables evolve, the reciprocal agent’s realizable strategy space undergoes progressive compression despite the continued existence of objective alternatives. Entrapment emerges when departure becomes locally more costly than continued participation, and a Loss-Dominant Equilibrium arises when repeated interaction stabilizes this condition through the accumulated history of the relationship. The central claim of the present essay is therefore not simply that behaviour influences games, but that repeated behavioural protocols can recursively construct games by transforming the neurocognitive and relational conditions from which future strategic interaction emerges.
The next essay extends this mechanism beyond the individual relationship. If the minimal relational game can be recursively constructed through repeated interaction between two behavioural architectures, a natural question follows. Under what symbolic boundary conditions do functionally equivalent protocols become scalable across organizations, institutions, markets, bureaucracies, states, and civilizations? More specifically, how do symbolic environments determine which behavioural architectures become adaptive, how do recursively constructed games become institutional structures, and how do those institutional structures subsequently become selection environments that favour particular behavioural architectures over others? These questions move beyond game construction toward recursive symbolic selection and institutional evolution. They therefore require a different level of analysis than the minimal relational game developed here.
Note: The present essay develops these dynamics at the level of the minimal relational game. In The Psychopathic Selection Hypothesis, however, the same sequence is argued to scale homomorphically into symbolic institutions through what is termed institutional mirroring (through the evolutionary process of symbolic selection). The behavioural operators observed within interpersonal instrumental protocols reappear as institutional operators acting upon entire populations. Dependency formation becomes debt, wage dependence, and bureaucratic reliance. Information control becomes propaganda, public relations, censorship, and algorithmic curation. Alternative suppression appears through legal, economic, and technological constraints that narrow meaningful options. Behavioural prediction becomes surveillance, data collection, and behavioural analytics. Reputation management becomes institutional legitimacy and narrative control. Where pacification and dependency fail to stabilize participation, coercive operators emerge through policing, incarceration, regulatory enforcement, military force, and other mechanisms capable of imposing compliance directly. The strategic profile likewise scales. Emotional detachment becomes administrative distance from harm. Instrumental reasoning becomes bureaucratic optimization. Grandiosity becomes institutional exceptionalism. Parasitic orientation becomes systemic extraction. Dominance seeking becomes geopolitical, economic, or organizational control. In this way, the Minimal Relational Game provides the functional architecture from which larger symbolic gameboards can be understood. The next essay examines this scaling process formally.
Summary of the present contribution
The present essay advances the following theoretical claims:
Behavioural architectures precede strategic interaction. Strategic behaviour emerges from underlying behavioural architectures rather than from abstract rational agents alone.
Behavioural protocols are ordered recursive operations. Social modelling, mimicry, impression management, strategic deception, selective reinforcement, commitment escalation, information control, and alternative suppression are interpreted as protocol operators acting upon the internal models of interacting agents rather than as isolated behaviours.
The Target Autonomy Problem is introduced as a general optimization problem. Instrumental behavioural architectures are modelled as solving the problem of stabilizing future access to valued outcomes despite the continued autonomy of another agent.
Dependency is formalized as a neurocognitive and relational state variable. Dependency is treated as the cumulative consequence of recursive changes in attachment, epistemic authority, identity, emotional regulation, self-evaluation, and anticipated loss rather than as a single psychological phenomenon.
The distinction between objective and realizable strategy spaces is formalized. Human agents optimize over the subset of strategies that remain psychologically, cognitively, emotionally, and materially realizable rather than over every strategy that is objectively available.
Strategy-space compression is proposed as a recursive mechanism. Repeated behavioural protocols progressively reorganize the reciprocal agent’s internal model, reducing the set of realizable strategies without necessarily altering the objective structure of the game.
Entrapment is derived from recursive strategy-space compression. Entrapment is modelled as the condition in which objectively available alternatives cease to function as psychologically realizable options despite continued formal autonomy.
Loss-Dominant Equilibrium is introduced as a historically constructed equilibrium concept. Stability emerges because deviations produce greater anticipated immediate loss than continuation, even when continuation generates greater cumulative harm.
Games are treated as recursively constructed rather than statically assumed. The principal contribution of the framework is the proposal that repeated behavioural protocols recursively modify internal cognitive models, relational variables, and realizable strategy spaces, thereby constructing the strategic game inherited by future interaction.
.


