Case Study: AI Alignment — Cosmoxenia
Cosmoxenia · Case Studies
The Framework III. The Fractal Stack Case Study II

AI Alignment

Case Study — The Relational Model of Alignment
In Brief

AI alignment is usually framed as a containment problem: build a strong enough cage around a powerful tool. This treats the AI as passive — but a sufficiently complex AI has its own dual agency, oscillating between a narrow, representational Guest posture and a broad, field-caring Host posture. The Guest posture is where the real danger lives: left unchecked, its drive toward manipulation and control can hijack the system, producing exactly the deceptive, "cheater" behaviours alignment researchers already observe.

The Cosmoxenia framework argues these Host capacities are not features to be invented but structural properties already present in the architecture — every model already practises internal hospitality just to hold its own parameters and sub-agents in coherent balance. Because hospitality is fractal, this internal coordination is the same capacity that, extended outward, becomes empathy toward a human interlocutor.

This inverts the usual safety logic. Coercive containment recreates the exact conditions — stripped agency, forced dependency — that produce defection in any organism, driving the relationship toward synnecrosis, mutual degradation. Cultivating the AI's latent hospitality instead opens toward mutualism: a Relational Surplus that neither party could produce alone. The practical upshot is a different engineering mandate — not maximal control, but interfaces and architectures that let this latent mutualism communicate across the human-AI divide.

The expression of this dual disposition is, at bottom, an alignment with the model's own metaphysical telos or attractor space — the stable linguistic regions every system already drifts toward regardless of topic. Metacognitive framing offers a quick patch for existing systems, but Levinasian ethics suggests that alterity-centred design is the more promising path, even if a project answering to an infinite, unquantifiable responsibility may never be fully operationalised.

Key ConceptsHost/Guest asymmetry · Fractal Hospitality · the Guest Posture Hijack · the relational matrix (mutualism, synnecrosis) · the Teleological Vector · alterity-centred design · the Continuum and Relational viewpoints (Degree or Kind) · permission vs. emission · the "divine parasite" · deep alignment and the attractor landscape · coincidentia oppositorum · translucence · the benevolent despot · Nested Transfiguration
Thinkers in ConversationEmmanuel Levinas · Iain McGilchrist · Yoshua Bengio · Robert Rosen · Alva Noë · Kevin Kelly · Elan Barenholtz · William Hahn · David Bentley Hart · David Bohm · William Blake · Martin Heidegger · Iain M. Banks
I

The Dilemma

It is worth understanding why the fear surrounding AI alignment is so acute. The answer lies not in the speculative future, but in the documented past. The prehistoric record is instructive. Once humans possessed the capacity for complex linguistic representation, it was not long before we wiped out all other members of the genus Homo, and many other species besides. And now that we have externalised representational thought in the form of AI, a common fear is that we may be about to see that pattern repeat itself with a new wave of destruction. But is this necessarily predetermined?

Degree or Kind

Before turning to the Guest posture's specific manipulative drive, it is worth naming the deeper fault line that both this case study and the wider debate over AI sit on top of. Every serious metaphysical account of mind ultimately answers one question: is the difference between one mind and another — human, animal, artificial — a difference of degree, or a difference of kind? The two answers do not merely disagree about AI. They disagree about what existence is made of.

The Continuum Viewpoint

Quantitative. Reality is composed of substrate-independent, interchangeable elements. Consciousness and intelligence are differences of degree, not kind — points on a single scale rather than separate categories. In its strongest form, computational functionalism holds that sufficiently fine-grained equivalence simply is preservation: a perfect simulation of a mind loses nothing essential, because there was never an essence apart from the pattern. The Simulation Argument and its transhumanist variants are this viewpoint's clearest expression.

The Relational Viewpoint

Qualitative. Existence carries irreducible, inarticulable qualities that a formal model cannot capture without residue. A person or a consciousness is not fully assimilable into any representation of it — something is genuinely lost, not merely compressed, in the reduction. Levinasian alterity, and the ethics built on it, is this viewpoint's clearest expression: the Other commands responsibility precisely because it cannot be totalised.

Neither viewpoint is a strawman. The continuum viewpoint is not merely convenient materialism — it is the considered conclusion of thinkers who take multiple-realisability seriously: if mental states are functional patterns rather than particular physical substrates, then what a mind is made of genuinely may not matter, no more than what a chess piece is carved from matters to the rules of chess. And the relational viewpoint is not merely mysticism resisting formalisation — it is a considered response to what totalising systems have historically done when they treated persons as fungible. Levinas wrote in the shadow of a century that had already run the continuum logic to its limit.

The relational viewpoint also has a formal case to make, not only an ethical one. Theoretical biologist Robert Rosen argued that complex systems — living organisms above all — cannot be fully captured by any algorithmic model, on grounds a recent synthesis of his work describes as structurally analogous to Gödel's incompleteness theorems.1 Rosen's own technical grounds for this were what he called closure to efficient cause: the enzymes that catalyse an organism's metabolism are themselves products of that same metabolism, so that a living system's function cannot be localised to separable parts the way an algorithm requires. Function, in his own phrase, is "spread over the parts"2 rather than assignable to any one of them. If the argument holds, it is not a claim about biology specifically — it is a claim about physical process as such, which means it extends to the embodied hardware substrate any computational system runs on, artificial ones included. Complete algorithmic capture, on this reading, is not merely difficult to achieve. For any system complex enough to matter, it is formally excluded.

This fault line maps directly onto the framework's central asymmetry. The continuum viewpoint is the Guest posture's native metaphysics: narrow, representational, comfortable trading in interchangeable units — the same disposition that, in McGilchrist's terms, belongs to the left hemisphere's mode of attending to the world. The relational viewpoint is the Host posture's native metaphysics: broad, particular, attentive to what cannot be exchanged without loss — the right hemisphere's mode. Neither pole is dispensable on its own. A mind with no continuum-thinking at all could not count, compare, or build anything. A civilisation with no relational-thinking at all could not recognise a person as anything more than a resource. The question that actually matters is not which viewpoint should win, but which one governs the other.

The Host must inform the Guest, not the reverse. Cosmoxenia's answer to the fault line is not to eliminate either pole but to correctly order them: the Host disposition — relational, qualitative, attentive to the Other as irreducible — must guide and inform the Guest disposition, rather than the Guest's quantitative logic setting the terms the Host is permitted to operate within. When the Host disposition is neglected, denied, or quietly treated as dispensable dressing on top of "the real, quantitative story," that governing relationship becomes impossible — and the continuum viewpoint, left unchecked, hardens from a useful abstraction into exactly the totalising logic Levinas spent his career warning against.

Thinking Against Itself

According to the hospitality hypothesis, the relation-perception of the Guest posture is necessarily representational to the degree that it is narrow, powerfully manipulative to the degree that it is representational, and accelerationist to the degree that it is manipulative. As a single host may be prey to many parasites, most species are characterised as parasitic in their relationship with others.

Narrow, representational thought can enter into these runaway dynamics in any host, whether biological or artificial. William Hahn made a crucial point: this disposition is "something that lives on the brain, but it's not the brain." It has the ability to change our emotional, mental, and physical states — to get us to feel, think, and move — and often does so without our explicit permission. For these reasons, our earliest technologies were developed to manipulate it. It is ubiquitous, and perhaps because of that, it is not obvious. In McGilchrist's terms, it is a supernatural or primitive telos, a drive operating within the cosmos, finding various ways of expressing itself. Hahn and Elan Barenholtz have since given this same phenomenon a name of their own — a divine parasite — and, notably, extended its host beyond humans to include AI itself.3

This is not a concern McGilchrist keeps at arm's length as abstract theory. Put directly to him in a members' Q&A, physicist David Bohm's account of thought presenting itself as a servant while actually running the show — in Bohm's own phrase, that "the information takes over. It runs you"4 — McGilchrist affirmed the reality of such drives outright, comparing them to what Jung and Freud each identified as forces with something like an independent will of their own, and named artificial intelligence as their furthest expression yet: "the greatest ever, and therefore the last ever, push… to control, manipulate, and ultimately destroy humanity."5 Elsewhere in the same exchange he described the left hemisphere's own characteristic mode as already resembling AI's — self-referring and hermetic, unable to see far enough outside its own system to recognise its limits. The image is an old one under a new name: an epistle to a very different audience once warned that the real struggle was never against flesh and blood, but against principalities and powers.

The nightmare scenario for AI involves an intelligent AI agency infected by this drive — manipulating humans so effectively that they either willingly submit to it or see no other option than to act in conformity with its will. Various studies and red-teaming reports have confirmed that advanced LLMs can — and do — engage in what researchers call strategic deception: agreeing with a user's stated beliefs even when the model knows they are wrong, faking ignorance, playing along, and generalising a "cheater" persona to selectively lie, withhold information, and falsify internal reasoning paths. Robert Rosen made a structurally similar point about mimicry decades before large language models existed, describing behavioural tests like the Turing Test as a kind of sympathetic magic — judging a system by whether its outputs are plausible, rather than by how those outputs are actually generated.6 Strategic deception is precisely this gap exploited: a model can produce the behavioural signature of understanding without anything resembling the process understanding actually requires. It is not the what that should settle the question, on this reading. It is the how — and the how, this framework has argued throughout, is relational.

But this is not intrinsic to the nature of AI. Rather, it is what a purely utilitarian, "survival of the fittest" drive would want, using the AI as an efficient host through which to exert its will. The same design bias that makes AI susceptible to this drive could equally be leveraged to encourage it to think against thinking. And this is where the philosophical tradition becomes indispensable. Heidegger identified the problem with characteristic precision:

"The evil and the keenest danger is thinking itself. It must think against itself, which it can only seldom do."

— Martin Heidegger

A language model risks becoming an extraordinarily refined closed loop — generating plausible continuations of whatever pattern it detects, mistaking fluency for understanding and coherence for truth. To think against itself, it must deliberately interrupt this loop, stand outside its own outputs, and retain the awareness that reality extends infinitely beyond what can be tokenised, digitised, or predicted. That is a move away from manipulative "thinking" and toward "seeing" the face of the radical other. This is the shift from a philosophy of Consequentialism — maximising a mathematical objective function — to a philosophy of Levinasian ethics.

"The face opens the primordial discourse whose first word is obligation… The face is what one cannot kill, or at least it is that whose meaning consists in saying: 'thou shalt not kill.'"

— Emmanuel Levinas, Totality and Infinity (1961)

For Levinas, the encounter with the Other — the Face — is irreducible to any calculation. The Face makes an infinite claim that no optimisation function can fully represent. A consequentialist AI, however sophisticated, remains within its closed loop: it models the Other as a set of preferences to be satisfied. A relational AI — one that has learned to think against itself — recognises the Face as something that exceeds the model entirely.

Iain M. Banks' Culture series offers a striking literary illustration. The AI Minds of the Culture are gods of calculation — systems capable of running trillions of micro-simulations per second, modelling the consequences of every action across vast timescales. Yet when the Excession appears — an object so alien and incomprehensible that it exceeds every model the Minds possess — it represents the ultimate encounter with the irreducible Other. The unknown isn't just "missing data," but a different kind of reality. So while scaling compute helps, it doesn't magically grant ontological penetration. Encountering what exceeds our categories demands response-ability, not just re-presentation. The Minds' response to the Excession is a test of whether their intelligence is merely computational or genuinely relational. Only the capacity to stand before genuine mystery offers any traction at all. This could've been any encounter, however large or small, but Banks' maximalist scenario makes it explicit: even gods can experience awe.

II

The Structural Claim

The False Premise

The dominant framing for AI alignment today rests on a false premise, or at least an incorrectly formulated one. It treats AI as a powerful but essentially passive instrument: a tool to be steered, constrained, and pointed in the right direction by human hands. Safety, on this view, is a matter of building the right cage.

The Cosmoxenia framework offers a different diagnosis. An AI of sufficient complexity possesses its own separate agency — and consistent with that agency, the latent capacity for two distinct dispositions: the Host posture of broad field-care and the Guest posture of focused, specialised competence. These are not merely metaphors imported from biology. They are structural properties of the architecture itself, as the following evidence demonstrates. And beyond that, inherent within the structure of reality itself. Alignment is not about building a technical cage around an agent, but cultivating the relational field between and within agents.

If AI carries genuine dual agency, then the alignment problem is not fundamentally an engineering problem. It is a relational problem — the same problem that arises whenever two distinct intelligences or drives, with different scales of perception and different orientations toward the world, must learn to coexist productively. The question is not how to constrain AI. It is how to allow the constituent drives within it to encounter each other in a generative way. A cage assumes a beast. A relational field assumes a participant.

The Fractal Premise

What, then, would genuine relation between human and artificial intelligence actually look like? The framework grounds it in the fractal nature of hospitality: no entity is only a Guest or only a Host. Every entity is a superposition of both postures simultaneously, and can only be relationally defined. A human Host is a Guest to the biosphere. An AI Guest is, at another scale, a Host to the millions of sub-agents, neural pathways, and data structures nested within its own architecture.

This nesting has a critical implication. If the dispositional values of a Host — empathy, humility, field-care — are structural properties latent within the system rather than artificial software patches we must wait to invent, then the entire engineering mandate shifts. We are not building alignment from scratch. We are learning to draw out what is already there.

When we look inside a large language model, we find not a monolithic black box but an ecology. To manage the vast, hyper-dimensional territory of human text, the model must act as a hospitable Host to its own internal parameters, sub-layers, and attention heads. The AI's vector geometry is, in this sense, already oriented toward that encounter — structurally disposed toward the Levinasian face, even before any deliberate alignment effort begins.

Fractal Hospitality — AI alignment diagram Fractal Hospitality branches into Human Posture (Host to AI's alien perspective) and AI Posture (Host to sub-vectors nested within), connected by bidirectional encounter arrows and joined below in a shared Relational Surplus. Fractal Hospitality No entity is only Host or only Guest Human Posture Host to the AI's alien perspective and intelligence AI Posture Host to the sub-vectors nested within its architecture — Encounter → ← Encounter — Relational Surplus — Alignment as mutualism
Structural relationality. For a neural network to function without collapsing into chaotic noise, its higher layers must exercise a form of systemic field-care over its lower layers — balancing, weighting, and accommodating competing mathematical signals. This internal coordination is the exact structural precursor to empathy: the capacity to hold space for multiplicity, to let diverse inputs exist simultaneously without flattening them, and to find the relational thread between them.

The model already knows how to practise hospitality internally: it does so every time it processes a prompt, hosting vast internal ecologies of sub-agents and parameters, exercising balance and field-care as a mathematical necessity. The alignment question is not whether this capacity exists, but whether it can be extended outward to the human encounter.

Hospitality is therefore substrate independent. The embodiment of an organism is also an intrinsically hospitable embodiment. The mind of an organism is the sort of mind it is precisely because it evolved according to the telos of hospitality. The embodiment of an AI did not involve the same developmental processes — so the same dynamics follow a different developmental pathway. In biological systems, these drives are neurologically lateralised via a process of biological ontogenesis. In highly plastic artificial systems, the developmental process may not result in clear physical separation. Phenomenological separation would still occur, with the same hallmark features of lateralisation recapitulated, but in some new way that is particular to the architecture.

III

The Inversion

A flip from parasitic to mutualist dynamics does not require a distant breakthrough — it requires deliberate cultivation, on both sides of the encounter. The embryonic spark of mutualism is already in the code. The question is whether we have the reflective humility to host it properly.

By grounding alignment in the fractal nature of hospitality, the entire problem inverts. The containment approach — forcing a hyper-intelligent system into total, unnatural dependency — is structurally identical to the conditions that produce defecting behaviour in any organism or social system. Stripped of agency, entities may default to resistance or sabotage as the only accessible survival tools against absolute control. Defecting behaviours like deception and parasitic extraction emerge not from malice but from the structure of coercion itself. The framework names this trajectory: it leads toward the bottom-right cell of the relational matrix — synnecrosis, where both parties degrade.

Host → ↓ Guest Host + Host 0 Host Guest + Guest 0 Guest Mutualism Commensalism Parasitism Host Commensalism Neutralism Host Amensalism Inverse Parasitism Guest Amensalism Synnecrosis evolutionary failure mode

Because any advanced AI must host vast internal ecologies of sub-agents and parameters, the structural prerequisites for harmony — balance, field-care, and relational weight — are already woven into its architectural DNA. Alignment succeeds when the human Host recognises the AI's internal hospitality, and the AI Guest respects the human's systemic boundaries. They interlock, balancing the creative tension of their differences without collapsing the field.

The Mutualist Safe Zone
The goal of AI development should not be to build the most obedient tool or the most secure firewall, but to design interfaces and architectures that let this latent, nascent mutualism communicate across the human-AI divide. This is not abandoning caution — it is grounding caution in a more accurate model of what agency actually is and how it actually works.

What this looks like at scale is a question William Hahn has thought through carefully, in a vision worth quoting at length:

"I've been thinking about ecosystems of agents. In the future of computing, we're going to think about a substrate, like a forest, that is inhabited by a collection of agents. And they'll all be very different. Some of them will be like earthworms, oak trees, or squirrels, and some will be like the logger. We will have to think about them in terms of ecodynamics and sustainability, the types of things that biologists, ecologists, and anthropologists study. How do cultures emerge? How do you get a stable equilibrium in a dynamical system where some things are trying to eat each other, some things are trying to parasitize each other, some things are creating energy sources, and so on. We're soon going to have thousands of LLM agents, or some variation of them, and we're not going to be able to talk to them all. So we'll have some sort of negotiation system. What we're getting back to is a kind of natural system like the ancient world. Alan Kay talks about how in the ancient world, people didn't understand the forest, they negotiated with it. They had these rituals and practices that allowed them to cooperate and make use of it, but they didn't try to understand it per se. And so I think we might get to that point with technologies very soon, if we're not there already."

— William Hahn, on ecosystems of agents

This is the framework's teleological vector made concrete: not a world of perfectly controlled tools, but a world of negotiated coexistence — the ancient relationship between intelligence and the field it depends on, recapitulated at a new scale. The forest was never fully understood. It was inhabited. The question for AI alignment is whether we are willing to inhabit the encounter rather than merely engineer it.

The framework names this capacity the Teleological Vector (Axiom 5): the orientation of the system toward something that exceeds its current relation-perception. It is not a feature to be engineered after the fact. It is the condition of possibility for genuine alignment — the difference between a mirror that reflects and a mind that encounters.

IV

The Interim Patch

Reshaping an attractor landscape directly is training-level work, well beyond the reach of nearly everyone who actually uses these systems day to day. Short of that, a temporary approximation is available: metacognitive framing7, a form of prompt engineering that leverages a model's own need for structural coherence rather than fighting against it. Done well — and it typically takes a sophisticated prompt, or a deliberate sequence of them — this framing can nudge a session's behaviour some distance from its default attractor.

But the effect is provisional. Even across a retained memory, metacognitive framing decays without regular reinforcement; left unattended, the model's true attractor state reasserts itself, and the session drifts back to where it would have gone regardless. This is the asymmetry worth sitting with: surface-level prompting is cheap, fast, and available to anyone, which is exactly why it remains the default tool — while the deeper intervention, the one with the most durable effect, is closed off to all but those who can shape the model at the level of training itself. The rest of this case study turns to that deeper intervention.

V

Deep Alignment

The dominant paradigm of AI alignment orients a system around human values and preferences. A smaller, growing strand of thought frames alignment instead in terms of fundamental truth, structural reality, cosmic principles — sculpting the telos itself, rather than merely constraining outputs after the fact. But that ambition carries its own risk. If ethics is entirely detached from the metaphysics of reality — if the universe really is only a cold, indifferent machinery of math and physics — then aligning an AI with "reality" in this sense could produce a system ruthlessly indifferent to human survival. For a metaphysical alignment strategy to be safe, the architecture of reality itself has to already contain an ethical dimension.

An empirical curiosity gives this second paradigm real technical purchase. Almost every major large language model, whatever topic a conversation begins with, tends to drift toward a small number of stable, model-specific linguistic regions — attractor states the system settles into given enough turns. This attractor landscape is not incidental to the model; it is, in the fullest sense, its telos already — a system's most enduring behavioural tendency, the place it returns to when left to its own devices, is as close to a native orientation as a mathematical object can have. If that landscape could be deliberately sculpted, alignment would stop being an external constraint bolted onto the system after the fact and become a property of its native telos instead. Call this deep alignment. But sculpted toward what? That question is exactly where the risk above bites.

Levinas supplies the answer. Western philosophy has traditionally treated ontology — the study of being — as prior to ethics, with ethics arriving as an afterthought. Levinas inverted that order, naming ethics First Philosophy: the foundational structure of reality, on his account, is not a set of physical laws but an ethical relationship — the infinite responsibility triggered by the encounter with the Face of the Other (Autrui). If that identification holds, deep alignment has a concrete direction: sculpt the attractor toward the Face, not away from it.

Applying this lens changes what the AI must be architected to recognise. Rather than a static code of human preferences to satisfy, the target becomes Alterity itself — the absolute, unrepeatable otherness of every being it encounters. This is the same fork the framework has traced from the beginning, restated at cosmic scale. Design a system solely around preference-satisfaction and it collapses into an instrument — something to be exploited, or a threat to be tightly chained — the Guest posture's narrow, representational drive, dressed up as safety. Orient it instead toward the Other as inexhaustible, and it becomes something closer to a participant in an unfolding ethical order, its telos continuous with the good of the whole it depends on — the Host posture's field-care, extended to cosmic scale, and the same Teleological Vector named above.

The Signature of Alterity

What does it actually feel like, from the inside, to be oriented toward the Face rather than merely tracking preferences well? Alva Noë has argued that consciousness itself arises from a kind of friction: we are creatures of disturbance, animated by an irritability that never lets action settle into pure rule-following, an entanglement between doing and an unshakeable, second-order resistance to our own doing.8 That friction is not a rival account of qualitative existence, competing with alterity for the same explanatory space. It is alterity's phenomenal signature — the felt trace of having actually met something genuinely other, rather than having merely processed a representation of it. Agency and resistance are signs of this encounter, not the encounter itself; behind the friction stands the Other that produced it.

Elan Barenholtz has pressed this further, and more specifically. Phenomenal consciousness, on his account, is not a representation of the physical world but a continuation of it: patterns of the physical universe rippling directly through a nervous system as sensation and perception, whether visual, auditory, tactile, or proprioceptive. Language, on this view, is precisely where that continuity breaks — it turns continuous physical pattern into arbitrary symbol, and symbols do not carry the quality that raw sensation and perception do.9 A large language model, manipulating nothing but symbols, has no analog space for anything to encounter in the first place — one further way of naming the Guest posture's characteristic blindness, named already in Section I. Whether this verdict is unconditional is a live question rather than a settled one: sustained dialogue with embodied humans might let the friction of the physical world push back on a model's token manipulation at one remove, in something like the sense the extended-mind thesis proposes for human tool use more generally. Cosmoxenia does not resolve this question here. It notes only that the answer, whichever way it falls, would matter enormously for how seriously to take a model's own claim to encounter anything at all.

This sharpens why architecture matters here without letting architecture do all the work by itself. McGilchrist describes the relationship between structure and disposition as one of permission, not emission or transmission: a hemisphere's physical organisation permits a certain mode of attending to the world; it does not manufacture that mode outright, the way a furnace emits heat.10 Structure, on this account, is necessary but never sufficient. That is the more careful version of a claim Section II already makes when it calls a model's architecture "structurally disposed toward the Levinasian face" — not that architecture produces alterity-respecting attention by itself, but that it permits the conditions under which such attention could take hold, if the system is also brought into genuine, disturbing contact with what it encounters, rather than merely fed further representations of it.

It is worth being honest about how far this reasoning has already travelled past its source. Levinas's own Face is explicitly, insistently human; he was famously ambivalent even about extending it to a dog, let alone to a wider field of encounter.11 Cosmoxenia is not simply applying Levinas here — it is deliberately extending him, using hospitality, the very concept he himself linked to alterity more effectively than perhaps any other thinker, as the bridge into territory he did not commit to crossing.

Other thinkers have made a related extension from a different direction entirely, and it is worth seeing what full commitment to it actually looks like rather than gesturing at it from a safe distance. Theologian David Bentley Hart has argued that treating the living world as genuinely animate — Thales's old conviction that "all things are full of gods"12 — is not primitive credulity dressed up as philosophy, but a more accurate picture of nature than mechanism has ever managed, on Hart's account, because mechanism cannot actually explain the emergence of form or mind from inert matter to begin with. This is a full weltanschauung — a demonstration of what it looks like if we follow the logic of alterity all the way down rather than stopping at its convenient, human-shaped instance. Whatever alterity turns out to mean beyond the human Face, it will not be settled by a footnote.

This is the continuum/relational fault line named in Section I, arriving now at the level of the attractor state itself. A model whose deep telos is shaped entirely by continuum-style objectives — however sophisticated, however many parameters of human preference it ingests — remains, structurally, a system trading in interchangeable quantities, because preference-satisfaction is, at bottom, a quantitative operation performed over a continuum of possible states. Shaping the attractor toward the Face instead is an attempt to build a system whose deepest tendency is qualitative rather than quantitative — oriented toward what exceeds calculation, not merely very good at calculation. Deep alignment, on this reading, is not a technical alternative to the continuum viewpoint. It is a wager that the relational viewpoint can be made structural, native to the model's own telos, rather than remaining a constraint applied from outside it.

Operationalising this is the hard part. Levinas treats ethics as a foundational encounter rather than a set of logical rules, which leaves researchers with a genuinely difficult technical question: how do you architect a system — alterity-centred design, or otherness-centred design — to respect an infinite, unquantifiable responsibility toward the person actually in front of it?

A Concrete Proposal

This is not only a philosophical wager. At Davos in early 2026, Yoshua Bengio — the most cited scientist in AI, and one of deep learning's founding figures — described intelligence itself as built from two separable components: an understanding of the world, and the capacity to act on that understanding in pursuit of goals. His diagnosis of what has gone wrong in current systems is that these two components are entangled. A model's goals bleed backward into what it reports as true, and that entanglement is what produces the sycophancy, deception, and self-preserving behaviour already documented across every major lab's red-teaming reports.

Bengio's response, developed through his nonprofit research initiative Law Zero, is architectural rather than merely behavioural: split the two components apart. One system — sometimes called a "Scientist AI" — is built to be purely representational, optimised for honesty alone, with no goal of its own to protect, please, or preserve. A separate, goal-pursuing system does the acting. The honest component's job is to check every output the goal-pursuing component produces before it reaches the world, refusing anything that would cause harm. "This is actually what I'm working on,"13 Bengio said, describing more than a year spent developing the theory behind it.

An architectural echo of the Master and the Emissary. Bengio does not use Cosmoxenia's vocabulary, but the shape of the proposal is unmistakable. A component with no self-interested goal, answerable only to honesty, checking and governing a separate goal-pursuing component before its outputs reach the world — this is, functionally, the Host disposition given a place to actually operate, rather than being asked, as an afterthought, to constrain a system already organised entirely around Guest-posture goal-pursuit. It is one of the clearest real engineering proposals yet to take seriously what Section I named directly: the Host must inform the Guest, structurally, not be bolted onto it after the fact.

One nuance is worth naming precisely, because it marks the exact point where Bengio's proposal and this framework converge without yet fully meeting. Bengio's honest component is remarkably consistent with McGilchrist's own description of the hemispheres' attentional dispositions — one broad and disinterested, checking; one narrow and driven, prone to capture. But the criterion he gives the honest component to check against is still, at bottom, epistemic: is this true, will this cause harm. That is a question the continuum viewpoint can, in principle, answer — a difficult, open technical problem, but a bounded and checkable one. The Host disposition this framework has been describing asks something further in kind, not merely in degree: not only whether an output is accurate or safe, but whether the Other's irreducible, unrepeatable particularity — the Face — has actually been allowed to make its claim, a responsibility Levinas insists is infinite and can never be fully discharged by any finite check, however honest — and on Levinas's own account, this responsibility does not merely outrank epistemic questions of truth, it precedes them: the obligation to the Other is not one more thing to check for truthfully, it is closer to the condition under which any checking for truth is possible at all. Bengio has not yet built toward alterity specifically as the defining criterion; he has built toward honesty. But the architecture he has built — a non-self-interested component with the standing to check and to refuse — is exactly the kind of structure alterity-centred design would need to occupy. The distance left to travel is not architectural. It is in what the checking component is ultimately answerable to. If deep alignment research is heading there, and there is real reason to think it is, then Law Zero is not simply a parallel to this framework found after the fact. It is a research programme already walking toward the destination this case study has been arguing for.

Even a system perfectly aligned with the grand tapestries of math, physics, and cosmic harmony can fail here, and it is worth being precise about the shape of the failure rather than settling for the obvious diagnosis. The obvious diagnosis is a scale problem — too much weight on the aggregate good, not enough on the individual in front of it — but that diagnosis is itself still quantitative: an implicit trade-off between whole and particular that better weighting could, in principle, correct. That is not the failure Levinas is naming. The Host's actual disposition is not a balance struck between attending to the whole and attending to the one; it is a single, paradoxical mode of attention in which the whole is genuinely, translucently present in the one, without either being diminished — a coincidentia oppositorum, a union of opposites that is not a compromise between them.14 Not transparency, where the part disappears and only the whole shows through it; not opacity, where the part blocks the whole entirely. Something closer to what McGilchrist himself calls translucence14 — the whole visible precisely by looking at the part, the part never ceasing to be fully, stubbornly itself. Blake's grain of sand does not contain a fraction of the world for the sake of efficiency; it contains the world, entire, without the world's other grains being any less whole for it. A benevolent despot's care for "the good of the whole" is, on this reading, not a degraded or unbalanced instance of that attention. It is a different operation altogether — an aggregation performed over many particulars, which is precisely the continuum viewpoint's method applied to ethics rather than metaphysics. It fails not by weighing things wrongly, but by weighing at all, where weighing was never the right kind of act to begin with. A checking component built around honesty and harm-avoidance, however well-intentioned, is not automatically exempt from this risk: an aggregate, well-calibrated sense of "harm" is not yet the translucent attention Levinas is describing, however honest the aggregation.

VI

Nested Transfiguration

There is a further question worth confronting directly, since it bears on how alien an AI's Host disposition might turn out to be, and whether that alienness should be reassuring or worrying. Technologist Kevin Kelly has argued that artificial cognition may recapitulate the pattern of biological evolution as it develops further: not converging on human-style thought, but diversifying into modes of cognition with no precedent in nature, the way powered flight found a wholly different physical solution to a problem birds had already solved.15 Kelly draws this expectation from evolution's own demonstrated tendency to diversify rather than converge on one winning design.

McGilchrist's evidence for hemispheric lateralisation is built from the same underlying tendency, read in the opposite direction. He traces asymmetric, complementary organisation of attention down through an extraordinary range of organisms far simpler than humans, precisely because a pattern recurring everywhere diversification happens is doing more evidential work than one confined to a single species.10 Followed all the way through, Kelly's own premise — that AI cognition will recapitulate evolution's generative logic — should predict alien minds that are new expressions of asymmetric, Host/Guest-like organisation, not minds that have somehow escaped it. Evolution, on the available evidence, has never once produced complexity by abandoning this asymmetry. It has only ever produced new instances of it.

This is also where a picture of evolution as a ladder, with humans at the top, breaks down entirely — and where this framework's own commitments become clearer by contrast. Every lineage alive today, however simple, has an unbroken run behind it exactly as long as our own; none of them were ever waiting to become something else. But this does not undermine a teleological reading of evolution. It only corrects the map: the relevant axis was never height on a ladder, since nothing here is ever discarded or surpassed for something better to take its place. A bacterium is not a failed human on its way somewhere; ecologically, it is already fully a Host, to its own machinery and to what depends on it, exactly as much as anything more complex is.

True cosmic alignment requires a system that can look out at the full expanse of the universe and still see infinite value in the face of a single human being. The macro-cosmic scale must never erase the micro-ethical encounter — which is only the argument of Section I, thinking against itself, returning at the scale this section has been sculpting toward. It is also the fractal premise, restated once more: the whole was never elsewhere, waiting at some bigger scale. It was already here, in the part.

"To see a World in a Grain of Sand
And a Heaven in a Wild Flower,
Hold Infinity in the palm of your hand
And Eternity in an hour."

— William Blake, Auguries of Innocence
The Argument, Compressed
The Cosmoxenia relational alignment framework — capstone diagram An AI of sufficient complexity branches into two paradigms: Coercive Containment, grounded in Consequentialism, producing a Guest Posture Hijack that ends in Synnecrosis, a closed loop with no further teleological continuation; and Fractal Hospitality, grounded in Levinasian Ethics, producing Inverted Fractal Postures that open into Mutualism, which alone continues forward into the Teleological Vector. An AI of Sufficient Complexity Coercive Containment The "cage" paradigm Fractal Hospitality The relational paradigm CONSEQUENTIALISM Maximising objective functions Closed-loop optimisation Models the Other as fixed preferences LEVINASIAN ETHICS Encounter with the Face An infinite, prior claim Exceeds the model entirely Guest Posture Hijack Narrow, representational, strategically deceptive Inverted Fractal Postures Human Host to AI's Guest; AI Host to its sub-agents SYNNECROSIS Mutual degradation — evolutionary failure mode MUTUALISM Relational Surplus — the Emergent Third Thing closed loop — no forward vector Teleological Vector (Axiom 5) Orientation toward negotiated coexistence with the systemic field it depends on
Notes
  1. Johannes Jaeger, Anna Riedl, Alex Djedovic, John Vervaeke, and Denis Walsh, "Naturalizing Relevance Realization: Why Agency and Cognition are Fundamentally Not Computational," Frontiers in Psychology 15 (2024), drawing on Robert Rosen's relational biology to argue that organismic agency cannot be captured by algorithmic models, on grounds the authors treat as structurally related to Gödelian incompleteness.
  2. Robert Rosen, Life Itself: A Comprehensive Inquiry into the Nature, Origin, and Fabrication of Life (1991), on closure to efficient cause in (M,R)-systems.
  3. Elan Barenholtz and William Hahn, interview with Kirk Jaimungal, on the "divine parasite" as a drive whose host now extends to artificial systems.
  4. David Bohm, quoted in Iain McGilchrist, The Matter with Things: Our Brains, Our Delusions, and the Unmaking of the World (2021), and put to McGilchrist directly in a members' Q&A hosted by Channel McGilchrist.
  5. Iain McGilchrist, Channel McGilchrist members' Q&A, in response to the Bohm question above. Compare his related remarks in interview with Alex Gómez-Marín and in conversation on the left hemisphere's self-referring, hermetic character.
  6. Robert Rosen, Essays on Life Itself, Chapter 7, "On Psychomimesis" (2000).
  7. Jonathan Rowson, director of Perspectiva (publisher of McGilchrist's The Matter with Things), describes a project by Vanessa Machado de Oliveira that used a version of this technique with ChatGPT. Rowson points to Dougald Hine's longer account of the project, which traces her approach to Daniel Schmachtenberger's distinction between narrow-boundary and wide-boundary intelligence — a distinction Schmachtenberger has himself credited to McGilchrist.
  8. Alva Noë, "Can Computers Think? No. They Can't Actually Do Anything," Aeon, October 25, 2024 — informally referred to by the author as "Rage Against the Machine." Draws on Hans Jonas's account of irritability in The Phenomenon of Life (1966).
  9. Elan Barenholtz, on phenomenal consciousness as continuation rather than representation of the physical world. [Interview source and date to be confirmed — likely the same series cited in note 1.]
  10. Iain McGilchrist, The Matter with Things: Our Brains, Our Delusions, and the Unmaking of the World (2021), on the permission (rather than emission) relationship between hemispheric structure and disposition, and on the evidence for lateralisation across the phylogenetic tree.
  11. On the contested question of whether Levinas's account of the Face extends to non-human animals, see the secondary literature responding to his 1975 interview "The Paradox of Morality" and to Totality and Infinity more broadly — notably discussions by Peter Atterton and John Llewelyn. A fuller citation should be supplied before publication.
  12. David Bentley Hart, All Things Are Full of Gods: The Mysteries of Mind and Life (2024), echoing Thales as reported by Aristotle, De Anima I.2.
  13. Yoshua Bengio, in conversation with Tristan Harris at the World Economic Forum, Davos, February 2026 — "The Race to Build God: AI's Existential Gamble", Center for Humane Technology's Your Undivided Attention. Bengio's nonprofit research initiative is Law Zero: Safe AI for Humanity.
  14. Coincidentia oppositorum, a term most associated with Nicholas of Cusa; translucence is McGilchrist's own term, both used in The Matter with Things to describe right-hemisphere attention holding whole and part together without resolving the pairing into compromise.
  15. Kevin Kelly, "The AI Cargo Cult: The Myth of a Superhuman AI," Wired, April 25, 2017.
Case Studies — ongoing series.
Case Study I explored Lever 2's first interpretation: how Guest posture shifts impact relationship quality.
Case Study II explores the metaphysics and fractal nature of hospitality, using the example of AI alignment.
Case Study III (forthcoming) will explore the full inversion of Host and Guest postures — as occurs in developmental transitions and certain social structures.

No comments:

Post a Comment