Small studio generative tools pipeline dev blog cover
Back to Dev Blog
engineering toolchain studio AI

Three People, One AI Pipeline: How We Build Lumenfall

Three people cannot hand-author an open world. The scale gap between team size and content ambition is the fundamental problem we set out to solve. This post describes the generative pipeline we assembled to close that gap, the infrastructure choices we made, and where the toolchain is honest about its limitations.

The Scale Problem Is a Content Problem

When we started building Lumenfall, the obvious question was: what does a three-person studio cut? The standard answer is scope. Make a smaller game. Tighten the map, reduce the NPC count, write fewer story branches. That answer is correct for most games. It is wrong for Lumenfall specifically, because the value proposition of Lumenfall is a world too varied for any player to exhaust. A smaller scope game with the same premise is incoherent.

The generative pipeline exists to make a different tradeoff: instead of cutting content scope, we cut per-unit authoring work to near zero. The terrain, characters, and story hooks that populate a Lumenfall world are assembled from authored primitives, not authored at the object level. We designed the grammar rules and the trait vocabularies and the narrative archetypes. We did not design each mountain, each NPC, each story event.

This is a real trade. The authored primitives constrain what the pipeline produces. We can express variety within the vocabulary we built. We cannot express things outside it. The pipeline produces a large combinatorial space, not an infinite creative space. We think the combinatorial space we have is large enough to be meaningfully infinite from a player's perspective. That hypothesis is what the closed alpha was testing.

Inference Infrastructure: What We Run and Why

The generation pipeline has two distinct computational profiles. Terrain grammar evaluation is CPU-intensive but does not require large memory or GPU time. It is deterministic logic running on structured data. Character and dialogue synthesis is different: it uses a locally-hosted language model for the personality profile generation step described in an earlier post, which requires GPU memory and has different latency characteristics.

We run the terrain grammar on the CPU using a multi-threaded zone evaluator that Tariq wrote early in development. It processes the world grid in parallel, resolving each zone's internal structure after the adjacency negotiation pass completes. On our development machines, full world generation runs in about 1.8 seconds from seed to complete zone data. That is fast enough for real-time re-generation on run start but not fast enough for in-play on-the-fly world expansion, which is why we pre-generate adjacent zones during active play rather than generating them at the moment of need.

For the language model step, we are running a quantized 8B parameter model on a dedicated GPU. The model runs as a background service that receives zone context prompts via a local socket and returns structured personality profile data. The service architecture means the model stays loaded in GPU memory between requests rather than reloading per zone, which is what makes the per-zone generation time short enough to run ahead of the player's movement.

The third infrastructure component is asset composition. Visual terrain assets are assembled from a vocabulary of modular pieces: rock formations, vegetation archetypes, geological detail elements. The composition system reads the zone type and feature tags and builds an asset configuration. This step is GPU-assisted for the LOD computation but is otherwise a lookup-and-placement operation rather than a generative one. We have about 2,400 unique asset pieces in the current build, which is a number a three-person studio can maintain because no asset needs to work as a standalone world element. Each one is a component in a larger assembly.

The Grammar Editor: Our Main Development Tool

The most important tool in the pipeline is the grammar editor, which is a custom tool Riya built during the first few months of development. It is not a beautiful piece of software. It loads in a web browser, communicates with the grammar engine over a local HTTP server, and the UI was designed by an engineer for engineers. But it does exactly what we need: you can edit grammar rules, immediately generate a preview world from a test seed, and see the change in the output.

The feedback loop the editor provides is the reason the grammar has evolved to 120-plus rules without becoming unmanageable. When a rule change has an unintended effect, the preview shows it immediately. When a new rule combination produces an interesting formation, you see it in the preview before it reaches a full game build. The editor handles the primary workflow of terrain design: write rule, test output, revise rule.

The editor does not handle cross-system integration. Testing how a grammar change interacts with the narrative attachment system or the character density model requires a full game build. That integration testing loop is slower and happens less frequently. Some of the more complicated bugs we have hit have been grammar changes that looked fine in the editor but broke something downstream that the editor cannot show. The gap between the editor's feedback and the full pipeline's behavior is a known cost of having a fast iteration tool for one layer.

What We Use Off the Shelf and Why

We built the spatial grammar engine, the trait pool system, the narrative archetype framework, and the grammar editor from scratch. We use off-the-shelf infrastructure for inference hosting, asset compression pipeline, session state serialization, and the audio system scaffolding.

The division is not arbitrary. We built custom systems where the design requirements were specific to Lumenfall's generative model and no available solution fit closely enough. We use off-the-shelf solutions where the problem is solved and the main requirement is that it works reliably. Spending engineering time reinventing an audio pipeline would be wrong. That time belongs on the generation systems that are actually the product.

The risk of this strategy is dependency. We are running a game that relies on infrastructure maintained by others. If the inference framework we use undergoes a breaking API change, or the compression pipeline drops support for a format we depend on, we have a maintenance problem. We have thought about this and decided it is the right tradeoff for a studio at our stage. We cannot build and maintain every layer of the stack. Managing dependencies is cheaper than building alternatives.

Where the Pipeline Honest About Its Limits

The pipeline produces a world that is varied in the dimensions we designed it to be varied in. It does not produce a world that is varied in the dimensions we did not design. If we did not build a rule for cultural artifacts, there are no cultural artifacts in the world. If the narrative archetype vocabulary does not include a tragedy archetype, no tragedy-shaped stories generate. The pipeline is bounded by its vocabulary.

We are also honest that the scale of what we can test is limited. We run generation on a few hundred distinct seeds per development cycle. The full space of valid seeds is enormous. Some configuration of inputs might produce something we have not seen and would not want to ship. The grammar rules are supposed to prevent this, but grammar rules have edge cases. The closed alpha was partly a test of whether real player-generated seed variation surfaced problems we had missed. The answer was yes, in a few cases. That is why we keep running playtests rather than assuming the rule set covers everything.

More from the Dev Blog

All posts
Alpha 04 terrain synthesis blog cover

Alpha 04: Terrain Synthesis Gets Smarter

Read post
Procedural terrain edge cases blog cover

When the Generator Surprises You: Terrain Edge Cases

Read post
Open world scale indie AI blog cover

Open-World Scale for an Indie Team: Why Generative AI Is the Only Path

Read post