LLMs for NPC personality generation dev blog cover with language model architecture diagram
Back to Dev Blog
AI NPC engineering lumenfall

Using Language Models to Generate NPC Personality Profiles

We spent about six weeks in early 2026 running experiments with a locally-hosted language model to generate NPC personality profiles from terrain context. The results were better than we expected and messier than we hoped. This post covers the architecture we settled on, the failure modes we hit before getting there, and what we would do differently starting fresh.

The Problem We Were Trying to Solve

Lumenfall's character synthesis system uses trait pool combinations to build NPC behavior profiles. The trait pool approach is solid: you get combinatorial variety without pre-scripting every character. But it generates traits, not personalities. A character with high suspicion, low sociability, and moderate ambition has a behavior profile. It does not have a coherent inner life that makes it feel grounded in the world it inhabits.

The gap we were trying to close was between "behavior profile that generates contextually appropriate responses" and "character that feels native to the specific terrain and region it exists in." A desert trader and a cliff-dwelling hermit can share the same trait profile. They should not feel like the same person. We wanted the local world context to do more work in shaping who a character is before any player interaction begins.

Language models are an obvious candidate for this kind of contextual synthesis. The question was whether we could run them at the scale and speed the generation pipeline required, on hardware a small studio actually has.

Infrastructure: Local Inference on Studio Hardware

The first constraint was that cloud API inference was not viable for per-character generation at world-load time. Round-trip latency on an API call would dominate zone loading even if cost were not a concern. We needed local inference.

We tested three configurations over the experiment period. The first was a 7B parameter quantized model running on a single consumer GPU. Generation speed was acceptable for batch processing but too slow for real-time zone loading. The second was a smaller 3B model with more aggressive quantization. Fast enough, but the personality profiles it produced lacked the contextual texture that made the larger model outputs useful. The third, and the one we are using now, is an approach we call deferred synthesis: the model generates personality profiles ahead of the player at zone pre-load time, not at the moment of encounter. This decouples generation latency from the player experience entirely.

Deferred synthesis means the model runs in a background thread while the player is inside an existing zone. By the time the player reaches the boundary of an adjacent zone, every NPC in that zone already has a complete personality profile. On our reference hardware, an 8B parameter model at 4-bit quantization processes a full zone's NPC complement in 14 to 18 seconds. Zone traversal time for a typical interior zone is 3 to 6 minutes of real play. The math works.

Prompt Architecture: Terrain Context as Input

The model receives a structured prompt that describes the terrain zone in constrained vocabulary. We deliberately avoided natural language zone descriptions because they produced inconsistent model behavior. Instead, the prompt format is a structured summary: zone type, adjacency types, elevation category, active weather state, and any named landmark features generated by the terrain grammar.

A sample prompt for a volcanic uplands zone adjacent to a thermal vent field looks roughly like this:

Zone type: volcanic uplands
Adjacent zones: thermal vent field (east), sedimentary basin (west)
Elevation: high plateau, 80th percentile
Weather state: ash drift, periodic
Terrain features: fractured basalt field, fumarole cluster (unnamed)
NPC role: resource specialist
Generate: personality profile, 3 primary traits, 2 behavioral tendencies, 1 knowledge domain

The output is structured too. We parse it into the trait pool format that the existing character synthesis system already understands. The LLM is not replacing the trait pool system; it is providing a contextually grounded starting configuration for it. The difference is that the trait selection is no longer random sampling from a flat pool. It is a directed selection shaped by where in the world the character lives.

What Went Wrong Before It Went Right

The first experiment version used free-form natural language prompts and asked the model to output a personality description. The outputs were creative and detailed and completely unusable. They used vocabulary outside our trait taxonomy, invented backstory elements we had not defined anywhere in the game world, and produced character concepts that were individually interesting but collectively incoherent as a population.

The lesson was straightforward: unconstrained language model output works against structured downstream systems. The model needs vocabulary constraints that match your data schema. We spent two weeks building the structured prompt format and output parser before any of the results were useful.

The second failure mode was terrain-personality non-sequiturs. Occasionally the model would generate a personality profile that was internally consistent but had no visible relationship to the terrain context. A character in a harsh volcanic zone with a cheerful, open, highly social personality profile. Not impossible in principle. But when the game generates populations this way, the aggregate effect is that the world feels disconnected from the people who live in it. We added a coherence filter that rejects profiles where fewer than two traits can be directly mapped to the terrain context vocabulary. Rejection rate is about 12 percent of first-pass outputs. The resampled second pass succeeds in over 95 percent of cases.

Results That Surprised Us

The most interesting output category was what we call emergent occupational specialization. Without any prompt instruction to produce this, the model consistently assigned knowledge domain traits that matched regional terrain features. Characters near the thermal vent field consistently generated volcanic geology knowledge domains. Characters near sedimentary basins generated mineral resource knowledge domains. The model was drawing on its training knowledge of how geography shapes human occupation, and that knowledge was producing plausible world-building without explicit instruction.

We are not taking credit for that. It is a property of how language models encode world knowledge. But it is a property we are actively building around because it is doing useful design work.

Cross-zone personality variation also increased noticeably in playtests after the LLM-generated profiles replaced the flat random trait sampling. Players described NPC populations in different regions as feeling distinct from each other in ways they could not always name precisely. The trait profiles are different in ways that are meaningful to the generation system but not necessarily legible as explicit differences to a player. The felt experience is still coherence and variety. That is what we want.

Where the System Still Falls Short

Factional and social network relationships between NPCs are not represented in the current profile generation at all. Each character is synthesized independently. Two characters who live in the same zone and share a knowledge domain could logically know each other, share history, belong to the same informal community. None of that exists in the current generation. The profiles are coherent individuals. They are not a coherent society.

Building relational context into the generation prompt is technically feasible and something Tariq has prototyped. The complication is ordering: to give character B information about character A's traits, character A must be synthesized first, which reintroduces sequential dependency into what is currently a parallelizable process. The tradeoff between relational richness and generation parallelism is a real design constraint we have not resolved. We are not saying it is impossible to solve. We are saying we have not solved it yet.

More from the Dev Blog

All posts
Alpha 04 terrain synthesis blog cover

Alpha 04: Terrain Synthesis Gets Smarter

Read post
50 players closed alpha playtesting blog cover

50 Players, 200 Hours: What Our Closed Alpha Taught Us

Read post
Character trait pool synthesis blog cover

Building Characters From Trait Pools, Not Pre-Written Scripts

Read post