Loading
Thesis Research
Stylized Deferred Rendering

This stylized deferred rendering pipeline runs on a custom DirectX 12 engine, inspired by Minecraft’s Iris and OptiFine shader ecosystem. It implements a G-Buffer pipeline with SM6.6 bindless resources, a ShaderBundle system with hot-swappable shader packs, and post-processing stages for volumetric clouds, volumetric lighting, screen-space reflections with three-layer blending, SSAO, bloom via compute-shader mipmaps, Lottes 2016 tonemapping with BSL-style color grading, water refraction, and sky glare. The engine layer handles chunk batch rendering with frustum culling, multi-frame in-flight GPU pipelining, and dirty-tracked root signature binding to reduce draw-call overhead.

0:00
/0:00
Demonstration of stylized rendering and in-app shader pack switching

Deferred rendering pipeline

The rendering pipeline uses a four-layer architecture: the application layer (Game) drives scene logic, the integration layer (RendererSubsystem) exposes public APIs, the core system layer manages DirectX 12 state and PSO caching, and the resource layer manages GPU memory through RAII D12Resource wrappers. All textures and buffers register into a single global descriptor heap with a 1,000,000 descriptor capacity and are accessed through SM6.6 bindless indexing. This avoids per-pass root signature slot bindings, allowing shaders to reference any resource directly through ResourceDescriptorHeap[index]. The application side drives execution through discrete RenderPass classes (Shadow, Terrain, TerrainCutout, TerrainTranslucent, SkyBasic, SkyTextured, Cloud, Deferred, Composite, and Final), each encapsulating its shader programs, texture bindings, and render states. Render targets, depth textures, and shader programs are resolved at runtime through a data-driven ShaderBundle configuration.

Several focused subsystems support the pipeline: DXCCompiler translates HLSL to DXIL with include graph resolution and comment-directive parsing, UniformManager manages 11 constant buffers through the Template Method pattern, PSOManager lazily caches pipeline state objects keyed by shader, render state, and render target formats, and provider classes (ColorTextureProvider, DepthTextureProvider, SamplerProvider) handle allocation, ping-pong flipping, depth copies, and dynamic sampler bindings.

graph TB
    subgraph APP["App Layer"]
        RP["RenderPasses"]
    end

    subgraph INT["Integration Layer"]
        RS["RendererSubsystem"]
        FQ["FullQuadsRenderer"]
    end

    subgraph CORE["Core Systems"]
        D3D["D3D12RenderSystem"]
        PSO["PSOManager"]
        UM["UniformManager"]
        DXC["DXCCompiler"]
    end

    subgraph BUNDLE["ShaderBundle System"]
        SBM["ShaderBundleManager"]
        FB["FallbackChain"]
        PROP["ShaderProperties"]
    end

    subgraph PROVIDER["Providers"]
        CTP["ColorTextureProvider"]
        DTP["DepthTextureProvider"]
        SP["SamplerProvider"]
    end

    subgraph RES["Resource Layer"]
        BRM["BindlessResourceManager"]
        GDH["GlobalDescriptorHeap"]
        BUF["D12Buffer"]
        TEX["D12Texture"]
    end

    RP --> RS
    RP --> FQ
    RS --> D3D
    RS --> PSO
    RS --> UM
    RS --> CTP
    RS --> DTP
    RS --> SP
    SBM --> DXC
    SBM --> FB
    SBM --> PROP
    PSO --> SBM
    PSO --> D3D
    CTP --> TEX
    DTP --> TEX
    UM --> BUF
    BRM --> GDH
    TEX --> BRM
    BUF --> BRM

Render target and depth providers

All render targets run through providers that implement the IRenderTargetProvider interface. Four concrete providers, ColorTextureProvider (colortex0 to 15), DepthTextureProvider (depthtex0 to 2), ShadowColorProvider (shadowcolor0 to 7), and ShadowTextureProvider (shadowtex0 to 1), own their GPU resources and share an API for creation, binding, clearing, and resizing. RenderTargetBinder aggregates the four providers behind a single facade, using pending and current state hashes to avoid redundant OMSetRenderTargets calls.

Color and shadow-color targets use dual-texture ping-ponging: each D12RenderTarget holds a main and alternate D12Texture, while a BufferFlipState<N> bitset tracks which buffer is the active write target for that slot. Calling Flip(index) swaps read and write assignments without a GPU copy. Depth textures are single-buffered. DepthTextureProvider snapshots depth at key stages using CopyDepth(), such as copying depthtex0 into depthtex1 before drawing translucent geometry. Every texture registers into the global bindless descriptor heap at creation. Each provider writes an IndexUniforms constant buffer mapping slot indices to bindless SRV handles, updated each frame with UpdateIndices() and uploaded to dedicated registers (b3 to b6) so shaders can sample any target by index.

graph TB
    RTB["RenderTargetBinder"]

    subgraph Providers
        CTP["ColorTextureProvider"]
        DTP["DepthTextureProvider"]
        SCP["ShadowColorProvider"]
        STP["ShadowTextureProvider"]
    end

    subgraph Resources
        DRT["D12RenderTarget"]
        DDT["D12DepthTexture"]
    end

    BFS["BufferFlipState"]
    IDX["IndexUniforms"]
    GDH["GlobalDescriptorHeap"]

    RTB --> CTP
    RTB --> DTP
    RTB --> SCP
    RTB --> STP

    CTP --> DRT
    SCP --> DRT
    DTP --> DDT
    STP --> DDT

    CTP --> BFS
    SCP --> BFS

    CTP --> IDX
    DTP --> IDX
    SCP --> IDX
    STP --> IDX

    DRT --> GDH
    DDT --> GDH
    IDX -.-> GDH

Vertex layout registration

Different passes require distinct vertex formats. A full-screen quad needs only position, UV, and color (PCU, 24 bytes), whereas terrain geometry packs normals, lightmap coordinates, block entity IDs, and mid-texture coordinates for atlas animation (TerrainVertex, 56 bytes). The engine manages these through a VertexLayoutRegistry. The abstract VertexLayout base class exposes GetInputElements() and GetStride(), concrete implementations (Vertex_PCU, Vertex_PCUTBN, TerrainVertex) register at startup, and the registry resolves const VertexLayout* pointers by name. Because this pointer is hashed into PSOKey, PSOManager caches a separate PSO for each shader-layout combination.

Each RenderPass declares its layout in BeginPass() with SetVertexLayout(). Terrain passes bind TerrainVertexLayout, while other passes use the default Vertex_PCUTBN. The renderer resets to the default at the start of every frame, so passes that omit SetVertexLayout() fall back predictably. Game-layer layouts register after engine startup to avoid engine dependencies on game types. TerrainVertexLayout also provides a MulticastDelegate event (OnBuildVertexLayout) that lets the ShaderBundle system inject material IDs into vertices when building meshes, without coupling the vertex type or chunk mesh to shader packs.

graph LR
    subgraph Registry["VertexLayoutRegistry"]
        VL["VertexLayout"]
        PCU["Vertex_PCU"]
        PCUTBN["Vertex_PCUTBN"]
        TV["TerrainVertex"]
    end

    VL --> PCU
    VL --> PCUTBN
    VL --> TV

    RP["RenderPass"] -->|SetVertexLayout| RS["RendererSubsystem"]
    RS -->|layout ptr in PSOKey| PSO["PSOManager"]
    PSO -->|GetInputElements| VL

Per-target blend configuration

Different render targets in the same pass often need separate blend configurations. A water pass may alpha-blend color into colortex0 while writing normals to colortex4 with blending disabled to preserve G-Buffer values. DirectX 12 supports this through IndependentBlendEnable, which the engine configures through two paths: shaders.properties directives and C++ RenderPass calls.

In shaders.properties, authors declare blend modes per program and per colortex slot. A directive like blend.gbuffers_water = SRC_ALPHA ONE_MINUS_SRC_ALPHA ONE ONE_MINUS_SRC_ALPHA sets the pass default, while a per-buffer override like blend.gbuffers_water.colortex4 = off disables blending for that slot. When loading a bundle, ShaderProperties parses these directives and InjectBlendDirectives() maps semantic colortex indices to physical slot indices through drawBuffers, storing them in ProgramDirectives. Render passes can also set blend modes directly by calling SetBlendConfig(config) or SetBlendConfig(config, rtIndex).

At draw time, RendererSubsystem packs the active blend state (global configuration plus up to 8 target overrides) into the PSOKey. PSOManager populates all 8 slots with the global configuration first, then applies overrides where isUndefined is false. This sentinel distinguishes explicit opaque settings from unconfigured slots, allowing unassigned targets to inherit the global default. Each pass resets its blend state to Opaque() after drawing to prevent settings from leaking into subsequent passes.

graph TB
    subgraph Config["Configuration Sources"]
        SP["shaders.properties"]
        CODE["RenderPass Code"]
    end

    subgraph Processing
        PARSE["ShaderProperties"]
        INJ["InjectBlendDirectives"]
        PD["ProgramDirectives"]
    end

    subgraph Runtime
        RSUB["RendererSubsystem"]
        PKEY["PSOKey"]
        PSOMGR["PSOManager"]
        D3D["D3D12_BLEND_DESC"]
    end

    SP --> PARSE --> INJ --> PD
    CODE --> RSUB
    PD --> RSUB
    RSUB --> PKEY --> PSOMGR --> D3D

DXC compiler

The engine compiles HLSL shaders at runtime through DXCCompiler, a wrapper around Microsoft’s DirectX Shader Compiler. Shaders target Shader Model 6.6 with HLSL 2021 syntax and 16-bit type support, producing DXIL bytecode for PSOManager. Because resource access is bindless, shader reflection (ID3D12ShaderReflection) is omitted. Resource indices are passed through root constants instead of reflected binding slots, reducing compile times and avoiding binding mismatches.

Before source text reaches DXC, IncludeGraph builds an include tree, IncludeProcessor flattens it into a single translation unit, and CommentDirectiveParser extracts render state from comments. DXC receives a self-contained string without #include statements.

Include graph

ShaderBundle HLSL files use #include for shared libraries such as lighting math, noise helpers, and uniform declarations. The engine resolves includes before compilation using a two-phase graph walk instead of DXC’s default file system include handler. First, IncludeGraph runs a breadth-first search from the entry file, discovering transitive dependencies and building a directed acyclic graph of FileNode objects. Each node stores a normalized virtual path and its child includes. Next, IncludeProcessor traverses the graph in depth-first order and concatenates files into one expanded source string, deduplicating shared headers. DXC then compiles this in-memory string without file system access.

Caching the include graph also supports incremental checks: when a shared header changes, the engine recompiles only dependent programs.

Comment directive parsing

Iris-style shaders encode pipeline state in HLSL comments rather than separate configuration files. CommentDirectiveParser scans the expanded pixel shader source and populates ProgramDirectives. Supported directives include:

  • /* RENDERTARGETS: 0,3,4 */ or /* DRAWBUFFERS:034 */ to specify color attachments, setting NumRenderTargets and RTVFormats[] in the PSO
  • /* DEPTH_TEST: LEQUAL */ and /* DEPTH_WRITE: true */ for depth-stencil state
  • /* CULLFACE: BACK */ for rasterizer culling
  • /* BLEND: ADD */ for blend operations

PSOManager reads these directives when constructing a PSOKey, creating and caching a PSO for each combination of draw buffers, depth mode, culling, and blend configuration.

graph LR
    subgraph Include["Include Resolution"]
        IG["IncludeGraph BFS"]
        IP["IncludeProcessor DFS"]
    end

    subgraph Directives["Directive Extraction"]
        CDP["CommentDirectiveParser"]
        PD["ProgramDirectives"]
    end

    HLSL["HLSL Source Files"] --> IG --> IP --> FLAT["Expanded Source"]
    FLAT --> DXC["DXCCompiler SM6.6"]
    FLAT --> CDP --> PD

    DXC --> DXIL["DXIL Bytecode"]
    PD --> PKEY["PSOKey"]
    DXIL --> PSOMGR["PSOManager"]
    PKEY --> PSOMGR
    PSOMGR --> PSO["ID3D12PipelineState"]

Built-in shader library

The engine includes a shader library with built-in programs, uniform declarations, and utility functions. Shaders can include "../include/core.hlsl" to bring in uniform constant buffers, vertex structures (VSInput/VSOutput), standard transforms, and math constants.

Uniform buffers are divided across two register spaces. Space 0 contains engine-managed buffers populated per frame or draw call: transform matrices (b7), camera parameters (b9), viewport dimensions (b10), per-object parameters (b1), and bindless index tables for color textures (b3), depth textures (b4), shadow colors (b5), shadow textures (b6), samplers (b8), and custom images (b2). Space 1 is reserved for game-level and bundle-defined uniforms matching Iris variables: world time (b1), fog parameters (b2), world info such as cloud height and ambient light (b3), rendering state such as rain strength and sky color (b8), and celestial data including sun angle and shadow light position (b9). The engine root signature binds space 0 through root CBVs, while space 1 binds through a descriptor table populated by the game layer, letting shader authors add custom buffers to space 1 without altering engine code.

All texture access is bindless. Instead of declaring explicit Texture2D : register(t*) bindings, include files store bindless SRV indices in uint4 arrays inside constant buffers, and shaders sample through ResourceDescriptorHeap[index]. Packing into uint4 arrays avoids the HLSL padding rule where scalar arrays pad each element to 16 bytes. Include files also supply helper functions: LinearizeDepth() and LinearToNDCDepth() in camera uniforms, CalculateFogFactor() and ApplyFog() in fog uniforms, and sampler aliases (linearSampler, pointSampler, shadowSampler, wrapSampler) in sampler uniforms.

engine/shaders/
  core/
    core.hlsl                    main entry, includes all uniforms
    gbuffers_basic.vs/ps.hlsl
    gbuffers_textured.vs/ps.hlsl
  include/
    camera_uniforms.hlsl         b9 space0, LinearizeDepth()
    matrices_uniforms.hlsl       b7 space0, all MVP/shadow matrices
    viewport_uniforms.hlsl       b10 space0, resolution and aspect
    color_texture_uniforms.hlsl  b3 space0, colortex0-15 indices
    depth_texture_uniforms.hlsl  b4 space0, depthtex0-2 indices
    shadow_color_uniforms.hlsl   b5 space0, shadowcolor0-7 indices
    shadow_texture_uniforms.hlsl b6 space0, shadowtex0-1 indices
    sampler_uniforms.hlsl        b8 space0, sampler aliases
    perobject_uniforms.hlsl      b1 space0, model matrix/color
    custom_image_uniforms.hlsl   b2 space0, custom texture indices
    common_uniforms.hlsl         b8 space1, renderStage/rain/sky
    worldtime_uniforms.hlsl      b1 space1, worldTime/moonPhase
    fog_uniforms.hlsl            b2 space1, fog color/density
    worldinfo_uniforms.hlsl      b3 space1, cloud height/ambient
    celestial_uniforms.hlsl      b9 space1, sun/moon/shadow pos
    developer_uniforms.hlsl      debug-only uniforms
  lib/
    fog.hlsl                     fog calculation utilities
    spaceConversion.hlsl         coordinate space transforms
  program/
    gbuffers_terrain.vs/ps.hlsl  terrain rendering
    gbuffers_water.vs/ps.hlsl    water surface
    shadow.vs/ps.hlsl            shadow map generation
    composite.vs/ps.hlsl         post-process compositing
    final.vs/ps.hlsl             final output to swapchain
    ...                          sky, cloud, debug programs

Compute shader and mipmap pipeline

The engine supports compute shaders within the same pipeline caching infrastructure used for graphics shaders. ShaderProgram and PSOManager handle compute PSO creation and caching, while UniformManager provides a PerDispatch update frequency for compute-specific constant buffers. The bindless root signature includes a ROOT_CBV_MIPGEN slot (b11) for mipmap generation uniforms.

MipmapGenerator produces GPU-driven mipmaps through compute shader dispatch. It provides four filter modes (Box, AlphaWeighted, SRGB, AlphaWeightedSRGB) compiled as shader variants during initialization through RendererEvents::OnPipelineReady. Generating each mip level dispatches ceil(w/8) x ceil(h/8) thread groups that read from the source mip SRV and write to the target mip UAV, with UAV barriers inserted between levels. Dispatches route to the compute queue when available and fall back to the graphics queue with cross-queue synchronization.

Render targets enable mipmap generation through colortexNMipmapEnabled directives in shaders.properties. CompositeRenderPass calls generateMipmapsForMarkedTargets() between composite sub-passes, allowing later passes such as composite4 bloom to sample hardware mipmaps created earlier. Atlas border extrusion prevents color bleeding at texture boundaries during downsampling.

Multi-frame in-flight rendering

The engine supports configurable multi-frame GPU pipelining (1 to 3 concurrent frames, defaulting to 2), allowing the CPU to record frame N while the GPU processes frame N-1. Frame-partitioned resource layouts use the active frame count to size ring buffers, descriptor tables, and per-frame allocations. FrameSlotAcquisitionResult tracks wait times and retirement statistics for profiling. Fence-based synchronization avoids resource hazards across frames and prevents CPU stalls during steady-state rendering.

Dirty-tracked root signature binding

GraphicsRootBinder tracks binding state to minimize redundant D3D12 calls. It maintains per-slot validity flags for the root signature, 15 engine CBV slots, and descriptor tables. Each Bind*IfDirty() call returns whether a rebind actually occurred, recording diagnostics such as bind counts, cache hits, and invalidations. Selective invalidation (InvalidateAll() versus InvalidateDescriptorTables()) provides cache control when pipeline state changes, reducing rebinding overhead on draw-heavy passes.

Per-pass uniform scope

The PerPass uniform scope isolates custom image bindings across composite sub-passes. SceneRenderPass tracks a pass scope that commits snapshots of dirty CustomImage bindings at pass boundaries, so textures bound in composite4 do not leak into composite1.

Chunk batch rendering

The chunk batch system groups voxel terrain into 4x4 chunk regions, each managing separate vertex and index buffer allocations across opaque, cutout, and translucent layers. ChunkBatchCollector tests region world bounds against the view frustum via ICullingVolumeProvider, producing a ChunkBatchCollection of visible items. ChunkBatchRenderer submits these batches using base vertex and start index offsets.

GPU buffer allocations use an arena manager that supports in-place updates and buffer relocation when a mesh expands. A PlayerCameraRig provides separate cameras for gameplay, rendering, debug inspection, and culling, making it possible to freeze the culling frustum while moving the debug camera freely. ChunkBachingRenderPass provides debug visualization with color-coded region wireframes (green for visible, red for culled, yellow for dirty, and magenta for rebuild errors).

graph TB
    subgraph Collection["Frustum Culling"]
        CAM["PlayerCameraRig"] --> FRUST["ICullingVolumeProvider"]
        FRUST --> COLL["ChunkBatchCollector"]
    end

    subgraph Regions["4x4 Chunk Regions"]
        R1["Region (Opaque)"]
        R2["Region (Cutout)"]
        R3["Region (Translucent)"]
    end

    COLL --> R1
    COLL --> R2
    COLL --> R3

    subgraph Submit["GPU Submission"]
        REND["ChunkBatchRenderer"]
        ARENA["Arena GPU Buffers"]
    end

    R1 --> REND
    R2 --> REND
    R3 --> REND
    REND --> ARENA

Application-side render pipeline

RendererSubsystem provides stateless APIs to bind shaders, set blend states, submit geometry, and upload uniforms, leaving pass ordering to the application layer. Game::RenderWorld() executes each RenderPass in a linear sequence, allowing the engine core to support forward, deferred, or hybrid layouts without changes.

Each pass derives from SceneRenderPass, an abstract base class with three lifecycle methods: Execute() (entry point), BeginPass() (state setup, shader, and target binding), and EndPass() (cleanup). The base class subscribes to OnBundleLoaded and OnBundleUnloaded events to handle shader reloads, so derived passes only need to update their cached ShaderProgram pointers. A utility class, RenderPassHelper, maps drawBuffers index lists to the typed render target references expected by UseProgram().

Game manages passes as unique_ptr<SceneRenderPass> instances and calls them in an Iris-compatible order within RenderWorld(): shadow generation, sky rendering into the G-Buffer, opaque terrain geometry, deferred lighting, translucent geometry (water and clouds), multi-stage compositing (SSR, volumetric light, tonemapping), and final output to the swapchain. Deferred lighting completes before translucent passes so transparent geometry blends over lit surfaces.

graph TB
    subgraph Shadow["1. Shadow"]
        S1["ShadowRenderPass"]
        S2["ShadowCompositeRenderPass"]
    end

    subgraph Sky["2. Sky"]
        SK1["SkyBasicRenderPass"]
        SK2["SkyTexturedRenderPass"]
    end

    subgraph GBuffer["3. Opaque G-Buffer"]
        T1["TerrainRenderPass"]
        T2["TerrainCutoutRenderPass"]
    end

    DEF["4. DeferredRenderPass (SSAO + Sky Glare)"]

    subgraph Trans["5. Translucent"]
        TT["TerrainTranslucentRenderPass (SSR + Refraction)"]
        CL["CloudRenderPass"]
    end

    subgraph Comp["6. CompositeRenderPass (Multi-Sub-Pass)"]
        C0["composite (SSR opaque)"]
        C1["composite1 (VL + Rainbow)"]
        MIP["Mipmap Generation"]
        C4["composite4 (Bloom Tile Atlas)"]
        C5["composite5 (Bloom + Tonemap + Color Grading)"]
    end

    FIN["7. FinalRenderPass (Underwater Distortion)"]
    CBR["8. ChunkBachingRenderPass"]
    DBG["DebugRenderPass"]

    S1 --> S2 --> SK1 --> SK2 --> T1 --> T2 --> DEF --> TT --> CL --> C0 --> C1 --> MIP --> C4 --> C5 --> FIN
    FIN --> CBR
    CBR -.-> DBG
Code/Game/Framework/RenderPass/
  SceneRenderPass.hpp/cpp           abstract base class (with PerPass scope)
  RenderPassHelper.hpp/cpp          static utilities
  WorldRenderingPhase.hpp           phase enum
  ConstantBuffer/                   POD uniform structs
  RenderShadow/                     shadow map generation
  RenderShadowComposite/            shadow post-process
  RenderSkyBasic/                   sky dome, void, and sky glare
  RenderSkyTextured/                sun, moon, stars
  RenderTerrain/                    opaque terrain (chunk batch integration)
  RenderTerrainCutout/              alpha-tested foliage
  RenderTerrainTranslucent/         water (SSR + refraction) and ice
  RenderCloud/                      cloud geometry
  RenderDeferred/                   deferred lighting (SSAO + sky glare)
  RenderComposite/                  multi-sub-pass: SSR, VL, mipmap, bloom, tonemap
  RenderFinal/                      output to swapchain (underwater distortion)
  RenderChunkBaching/               chunk batch debug visualization
  RenderDebug/                      debug overlays

The ShaderBundle system adapts the concept of Minecraft’s Iris and OptiFine shader packs. A ShaderBundle is a directory of HLSL programs, property files, fallback rules, and textures that defines scene rendering. The engine discovers bundles by scanning .enigma/shaderbundles/ at launch and supports runtime switching through an ImGui interface or code API.

The system uses a dual-bundle design. An internal engine bundle ships with the renderer as a permanent baseline. When a user bundle is active, GetProgram() resolves shaders through a three-level chain: active user-defined sub-bundles for quality variants, the bundle’s program/ directory using fallback_rule.json (such as falling back from gbuffers_clouds to gbuffers_textured and then gbuffers_basic), and finally the engine bundle to ensure every pass has a valid program. Bundle switches occur at frame boundaries to avoid deleting in-use render targets mid-frame.

Shader bundle system

A ShaderBundle directory under .enigma/shaderbundles/ follows a standardized layout:

  • program/: primary shader programs (.vs.hlsl and .ps.hlsl) for each pass
  • bundle/: named sub-bundles that override programs for profile or dimension presets
  • lib/: shared HLSL libraries (lighting, noise, shadow calculations, tonemapping) included by programs
  • include/: header declarations specific to the bundle
  • shaders.properties: render target formats, blend modes, and buffer configurations
  • block.properties: maps block IDs to material categories for MaterialIdMapper
  • textures/: custom texture assets with optional .enigmeta metadata files for sampling rules

On the engine side, ShaderBundle and UserDefinedBundle manage program references and fallback resolution. ShaderProperties and PackRenderTargetDirectives parse property files into configuration structs, MaterialIdMapper translates block IDs into material categories, and BundleTextureLoader loads auxiliary textures. ShaderBundleSubsystem coordinates discovery, lifecycle events, and settings persistence in shaderbundle.yml.

graph TB
    subgraph Integration
        SBS["ShaderBundleSubsystem"]
        CFG["Configuration"]
    end

    subgraph Core
        SB["ShaderBundle"]
        UDB["UserDefinedBundle"]
        PFC["ProgramFallbackChain"]
    end

    subgraph Config["Configuration Parsing"]
        SP["ShaderProperties"]
        RTD["PackRenderTargetDirectives"]
        PSD["PackShadowDirectives"]
        MIM["MaterialIdMapper"]
    end

    subgraph Assets["Asset Loading"]
        BTL["BundleTextureLoader"]
        EMP["EnigmetaParser"]
    end

    subgraph Helpers
        JH["JsonHelper"]
        FH["FileHelper"]
        SH["ScanHelper"]
    end

    SBS --> SB
    SBS --> CFG
    SB --> UDB
    SB --> PFC
    SB --> SP
    SB --> RTD
    SB --> PSD
    SB --> MIM
    SB --> BTL
    BTL --> EMP
    SB --> Helpers
EnigmaDefault/shaders/
  shaders.properties           RT formats, blend modes, buffer config
  block.properties             block ID to material mapping
  bundle.json                  bundle metadata
  program/
    gbuffers_terrain.vs/ps     terrain geometry
    gbuffers_water.vs/ps       water surface
    shadow.vs/ps               shadow map generation
    deferred1.vs/ps            deferred lighting
    composite.vs/ps            post-process pass 0 (SSR opaque)
    composite1.vs/ps           post-process pass 1 (VL + rainbow)
    composite4.vs/ps           bloom tile atlas generation
    composite5.vs/ps           bloom apply + tonemapping + color grading
    final.vs/ps                output to swapchain
    ...                        sky, cloud, debug programs
  bundle/
    mycustom_bundle_0/         sub-bundle variant (overrides program/)
      gbuffers_terrain.vs/ps
      composite.vs/ps
      ...
  lib/
    atmosphere.hlsl            sky and atmosphere math
    bloom.hlsl                 bloom tile atlas generation and reading
    clouds.hlsl                volumetric cloud ray marching
    common.hlsl                shared utilities (Luma, Pow2, Pow4)
    fog.hlsl                   fog calculations
    lighting.hlsl              diffuse and specular models
    noise.hlsl                 procedural noise functions
    pipelineSettings.hlsl      quality toggles and pipeline config
    rainbow.hlsl               procedural rainbow generation
    reflection.hlsl            three-layer reflection coordinator
    refraction.hlsl            water surface refraction
    shadow.hlsl                shadow sampling and bias
    skyGlare.hlsl              sun/moon atmospheric halo
    ssao.hlsl                  screen-space ambient occlusion
    ssr.hlsl                   screen-space reflection ray march
    tonemap.hlsl               Lottes 2016 HDR to LDR + color grading
    underwaterDistortion.hlsl  underwater screen-space distortion
    volumetricLight.hlsl       volumetric ray marching
    water.hlsl                 water surface utilities
  include/
    settings.hlsl              user-facing quality toggles
    ...                        bundle-specific declarations
  textures/
    cloud-water.png            custom texture asset
    cloud-water.png.enigmeta   sampling and format metadata

Material ID mapper

Voxel terrain quads share the same vertex and pixel shaders in each pass. Distinguishing water, foliage, stone, or emissive blocks within batched geometry requires per-vertex material identifiers.

MaterialIdMapper connects block.properties definitions to the vertex stream. Authors assign numeric IDs to namespaced block names (such as block.32000=simpleminer:water). When a bundle loads, the mapper parses these entries into an unordered_map<string, uint16_t> lookup table. It then registers a callback with TerrainVertexLayout::OnBuildVertexLayout, an event that fires as the voxel mesher generates quads. If a block matches a configured rule, OnBuildVertex() writes the material ID into the m_entityId field of the quad’s vertices. The pixel shader reads this integer from TEXCOORD2 (matching Iris’s mc_Entity semantic) to select material-specific shading.

This keeps the systems decoupled: the mesher processes geometry without material logic, the vertex layout stores an uninterpreted 16-bit integer, and the shader branches on that integer, while ID assignments remain in block.properties.

graph LR
    BP["block.properties"] --> MIM["MaterialIdMapper"]
    MIM --> EVT["OnBuildVertexLayout"]
    MESH["Voxel Mesher"] --> EVT
    EVT --> TV["TerrainVertex m_entityId"]
    TV --> PS["Pixel Shader TEXCOORD2"]

Lifecycle events

The engine manages bundle transitions through two event channels: a string-based EventSystem on the global event bus, and typed MulticastDelegate callbacks for direct, type-safe notification. Subsystems subscribe to these events to respond to bundle changes without circular dependencies.

EventMechanismTriggerTypical Subscriber
OnShaderBundleLoadedEventSystemAfter a user bundle finishes loadingRenderPasses (rebuild PSO cache)
OnShaderBundleUnloadedEventSystemBefore switching back to engine bundleRenderPasses (release user programs)
OnShaderBundlePropertiesModifiedEventSystemWhen shaders.properties is edited at runtimePSOManager (invalidate cached state)
OnShaderBundlePropertiesResetEventSystemWhen properties are reset to defaultsPSOManager (restore original config)
OnShaderBundleReloadEventSystemWhen a hot-reload is requestedShaderBundleSubsystem (recompile all)
OnBundleLoadedMulticastDelegateAfter bundle load completesMaterialIdMapper subscription setup
OnBundleUnloadedMulticastDelegateAfter bundle unloadMaterialIdMapper subscription teardown
OnBuildVertexLayoutMulticastDelegatePer quad during chunk meshingMaterialIdMapper (inject entity ID)

Delegates manage the MaterialIdMapper lifecycle directly. When a bundle loads, ShaderBundleSubsystem registers OnBuildVertex with TerrainVertexLayout::OnBuildVertexLayout and saves the DelegateHandle. On unload, it unregisters the handler with that handle, ensuring callbacks do not persist across bundle swaps.

Render target format and clear configuration

A deferred renderer stores distinct data formats across render targets: HDR color in colortex0 may require R16G16B16A16_FLOAT, normals in colortex2 use R8G8B8A8_SNORM, and a material mask in colortex3 uses R8G8B8A8_UNORM. Authors can declare target formats, clear behaviors, and clear colors directly in HLSL within rt_formats.hlsl, colocating configuration with shader code.

PackRenderTargetDirectives parses two syntax styles from rt_formats.hlsl:

  • Format directives reside inside /* */ comments because DXGI format tokens (such as R16G16B16A16_FLOAT) are not valid HLSL keywords. ConstDirectiveParser extracts them from raw source lines.
  • Clear flags and colors use standard HLSL const declarations (const bool colortex0Clear = true, const float4 colortex0ClearColor = float4(0,0,0,1)), parsed through constant evaluation.

Both forms support colortex (0 to 15), depthtex (0 to 2), shadowcolor (0 to 7), and shadowtex (0 to 1). When a bundle loads, the engine merges these directives with default YAML settings to produce a RenderTargetConfig per slot, which providers use to allocate textures and set frame clear actions.

// rt_formats.hlsl example (EnigmaDefault)

// Format directives (inside comments, not valid HLSL)
/*
const int colortex0Format = R16G16B16A16_FLOAT;
const int colortex1Format = R8G8B8A8_UNORM;
const int colortex2Format = R8G8B8A8_SNORM;
*/

// Clear control (valid HLSL const declarations)
const bool colortex0Clear = true;
const bool colortex6Clear = false;

// Clear color
const float4 colortex1ClearColor = float4(0.0, 0.0, 1.0, 1.0);

// Shadow RT configuration
const bool   shadowcolor0Clear      = true;
const float4 shadowcolor0ClearColor = float4(1.0, 1.0, 1.0, 1.0);
graph TB
    subgraph HLSL["rt_formats.hlsl"]
        FMT["Format Directives"]
        CLR["Clear / ClearColor"]
    end

    subgraph Parsing
        CDP["ConstDirectiveParser"]
        PRTD["PackRenderTargetDirectives"]
    end

    subgraph Output["RenderTargetConfig per slot"]
        CF["DXGI Format"]
        CA["Clear Action"]
        CC["Clear Color"]
    end

    FMT --> CDP
    CLR --> CDP
    CDP --> PRTD
    YAML["YAML Defaults"] --> PRTD
    PRTD --> CF
    PRTD --> CA
    PRTD --> CC
    CF --> PROV["Providers"]
    CA --> PROV
    CC --> PROV

Shader properties configuration

The shaders.properties file configures bundle-level pipeline options without C++ modifications:

  • Quality profiles define presets (such as POTATO through ULTRA) that set preprocessor macros and numeric values for shadow distance, reflection sample counts, and SSAO parameters.
  • Custom texture bindings map image files to numbered slots per program (for example texture.deferred.3=textures/cloud-water.png), accessible in HLSL through the bindless GetCustomImage() function.
  • Blend directives configure program and target blend operations.

At load time, ShaderProperties parses this file and distributes the settings: profiles set compiler macro definitions, texture bindings pass to BundleTextureLoader, and blend rules inject into ProgramDirectives for PSO generation.

Stylized shader bundle

EnigmaDefault is the primary ShaderBundle included with the project, inspired by the visual design of Complementary Reimagined for Minecraft. It prioritizes atmospheric depth, soft lighting, and stylized color grading over photorealism, using volumetric clouds, atmospheric scattering, multi-layer water reflections, and screen-space ambient occlusion. Every effect is implemented in HLSL inside lib/ and program/ and runs through the engine’s data-driven pipeline without dedicated C++ rendering paths.

Atmospheric scattering with warm sunset tones and depth fog
Atmospheric scattering at sunset, showing depth fog and horizon color blending

Shadow mapping and bias

The shadow system uses a single shadow map with nonlinear XY distortion instead of cascaded shadow maps. The distortion warps clip-space coordinates so regions near the camera receive higher texel density while distant areas are compressed, providing near-field detail in a single depth pass. The Z axis is compressed to 20% of its original range to expand depth precision across the frustum. The distortion factor is controlled by SHADOW_DISTANCE (configurable per quality profile, defaulting to 128 blocks).

Bias combines two techniques. Normal offset bias runs in world space, shifting the shadow sample position along the surface normal. The offset scales with distance to account for lower texel density and increases up to 2x at grazing angles where NdotL approaches zero, avoiding shadow acne without the light leaking that constant depth bias produces on thin geometry. A soft depth comparison window (factor of 256) normalized against the Z compression then sharpens transition boundaries before PCF filtering.

Shadow filtering supports six quality levels, from a single tap to 16-tap circular PCF. The PCF kernel uses Interleaved Gradient Noise (IGN) for screen-space dithering to prevent banding without temporal filtering. Kernel radius scales with distance from the camera, rain intensity (widening up to 3x in overcast conditions), and an edge fade band over the final 8 blocks before the shadow cutoff distance.

The shadow pass writes two auxiliary targets alongside depth: shadowcolor0 stores a dual-frequency water caustic pattern, and shadowcolor1 stores underwater volumetric light color with exponential distance falloff.

graph TB
    subgraph Shadow["ShadowRenderPass"]
        SVS["shadow.vs.hlsl"]
        SPS["shadow.ps.hlsl"]
    end

    subgraph Output["Shadow Output"]
        ST1["shadowtex1 depth"]
        SC0["shadowcolor0 caustics"]
        SC1["shadowcolor1 underwater VL"]
    end

    subgraph Deferred["DeferredRenderPass"]
        D1["deferred1.ps.hlsl"]
        LIT["lighting.hlsl"]
        SHAD["shadow.hlsl"]
    end

    SVS --> ST1
    SPS --> SC0
    SPS --> SC1

    ST1 --> SHAD
    SHAD --> LIT
    SC0 --> LIT
    SC1 --> LIT
    LIT --> D1
    D1 --> CT0["colortex0 lit scene"]

Volumetric clouds

The cloud system uses screen-space ray-marched volumetric clouds evaluated in the deferred lighting pass (deferred1). Cloud coverage is driven by a 2D texture lookup (cloud-water.png blue channel) with an 8th-power threshold for defined edges, instead of 3D noise textures. The shader includes self-shadowing via light-direction sampling, height-gradient shading (dark base, bright top), forward scattering from a half-Lambert view-sun dot product, and three sample counts (16, 32, and 48). Output cloud depth passes to composite stages so god rays do not illuminate through solid cloud bodies.

For a detailed breakdown of the ray march algorithm, texture-driven shape generation, self-shadow computation, and cloud lighting model, see the dedicated blog post: Volumetric Cloud in Deferred Rendering.

Volumetric clouds from above showing self-shadowing and height gradient shading
Volumetric clouds from above, with bright tops and self-shadowed undersides

Volumetric lighting

The volumetric light pass renders light shafts by marching rays from the camera toward each world-space fragment, accumulating lit steps by sampling the shadow map. Directional modulation through VdotL (view-to-light dot product) concentrates shafts toward the sun and moon. The implementation supports four sample tiers (12 to 50 samples for daytime), time-of-day color gradients from sunset warmth to moonlight, and noon dimming to prevent overexposure. A dual shadow test comparing shadowtex0 and shadowtex1 allows light shafts to pick up colored tinting through translucent surfaces using shadowcolor1. Underwater, ray marching is limited to 80 blocks and ignores fully lit samples so illumination comes exclusively from shafts entering through the surface.

For implementation details covering shadow sampling, directional modulation formulas, and underwater attenuation, see the dedicated blog post: Volumetric Light in Deferred Rendering.

Volumetric light shafts at sunrise viewed from a mountain
Volumetric light shafts at sunrise, showing directional modulation and warm sunset color grading
Volumetric light shafts at night with cool blue moonlight
Volumetric light shafts at night, with cool blue moonlight and shadow-driven ray attenuation

Screen-space and three-layer reflections

The water rendering system uses a three-layer reflection pipeline that blends near-field ray marching, mid-distance mirrored image reprojection, and far-field procedural sky across viewing distances.

Layer 1: SSR ray march

View-space ray marching uses exponential step increments and binary refinement, running inline in gbuffers_water to avoid allocating an intermediate composite target. The ray march tests against depthtex1 (opaque depth) to prevent self-intersections on the water surface, applying screen-border fading via pow(max(cdist.x, cdist.y), 50.0) and depth-proximity thresholds. Smoothness modulates the resulting alpha, allowing rougher water to fall through to lower reflection layers.

Layer 2: Mirrored image reprojection

When SSR rays exhaust their step limit but the reflected scene remains visible from another on-screen angle, this layer projects the reflected vector into clip space with vertical parallax stretching. It samples colortex0 at the reprojected coordinate, applies a power-8 edge fade, and adds exponential distance fog. Reflections oriented behind the camera are discarded by an angle check.

Layer 3: Procedural sky fallback

The final fallback generates an atmospheric gradient with zenith-to-horizon blending, sunset horizon coloring, and below-horizon darkening. Skylight attenuation prevents sky reflections from appearing in caves or indoor water.

The three layers blend in priority order: SSR overrides the mirrored image, which overrides the procedural sky. Uncovered regions fall through to lower layers based on alpha coverage. A GGX specular highlight adds sun and moon glints.

The water surface also includes three-layer normal distortion scrolling at varied rates, depth-based transparency against depthtex1, Fresnel blending with a cubic approximation ensuring 15% minimum reflectivity, and shoreline foam.

graph TB
    subgraph Layers["Three-Layer Reflection"]
        L1["Layer 1: SSR Ray March (near)"]
        L2["Layer 2: Mirrored Image (mid)"]
        L3["Layer 3: Procedural Sky (far)"]
    end

    L1 --> BLEND["Priority Blend"]
    L2 --> BLEND
    L3 --> BLEND
    GGX["GGX Specular"] --> BLEND
    BLEND --> FINAL["Final Reflection Color"]
    FRESNEL["Fresnel"] --> MIX["Reflection + Scene Mix"]
    FINAL --> MIX

For the complete implementation covering the ray march algorithm, normal distortion, depth-based culling, and underwater effects, see the dedicated blog post: Screen-Space Reflection Water in Deferred Rendering.

SSR water at noon with bright blue sky reflections
Screen-space reflections on water at noon, showing three-layer blending with depth-based transparency and Fresnel-driven reflection
SSR water at low view angle showing Fresnel-driven reflectivity
Screen-space reflections at a low view angle, demonstrating strong Fresnel reflectivity and mirrored image fallback blending

Water refraction

Above-water refraction uses noise-based UV offsets with three validation checks to prevent distortion artifacts. The shader samples a noise texture (customImage4) in world space animated by frameTimeCounter, so distortion tracks surface movement rather than camera orientation. Offset scale incorporates field of view compensation and inverse distance scaling to dampen distant distortion.

Three checks prevent edge bleeding:

  1. A material mask in colortex4 confirms the source pixel is water.
  2. Depth comparison between the water surface (depthtex0) and opaque geometry (depthtex1) scales down refraction in shallow areas.
  3. The warped offset position is re-checked against the water mask to ensure the sampled pixel remains within the water surface, preventing shoreline pixels from bleeding across boundaries.

Underwater effects

When the camera is submerged, underwater post-processing activates across multiple passes. The composite1 pass applies noise-based refraction, while the final pass adds sinusoidal screen-space distortion using a diagonal ripple animated by frameTimeCounter to produce a wavering view. Distance fog uses squared exponential falloff with wavelength-dependent color attenuation applied in gamma space before linearization. Colored volumetric light shafts enter through the surface using the dual-depth shadow tests (shadowtex0 and shadowtex1), tinted by shadowcolor1.

Underwater volumetric light shafts passing through the water surface
Underwater colored volumetric light shafts, with wavelength-dependent fog attenuation and screen-space distortion

Procedural rainbow

The composite1 pass renders a procedural rainbow when solar elevation falls between 0.1 and 0.25, centered roughly 42 degrees from the anti-solar point. The arc geometry uses a bell curve profile with configurable diameter, while spectral bands are evaluated through non-linear coordinate mapping across three phase-offset color channels. Rain intensity, cloud coverage, and solar elevation modulate opacity, with heavy dimming applied when the camera is underwater.

Procedural rainbow arc during light rain at sunset
Procedural rainbow arc at low solar elevation, with spectral color generation and rain-modulated visibility

Screen-space ambient occlusion

The deferred lighting pass (deferred1) evaluates ambient occlusion using depth-only Poisson disk sampling with bilateral depth tests, adapted from Complementary Reimagined. The algorithm works directly in screen space from depth samples without reconstructing G-Buffer normals, combining surface-angle metrics with distance thresholds to compute occlusion.

Sample offsets use a golden-ratio angular distribution with squared radial falloff, concentrating taps near the center of the fragment. Kernel radius scales with field of view and camera distance to match perspective foreshortening. The shader provides two presets: a fast mode with 4 bilateral samples at 0.4 scale, and a quality mode with 12 samples at 0.6 scale. Each step samples both +offset and -offset directions, doubling test coverage. The resulting factor is raised to an exponent (SSAO_IM) to control shadow intensity.

SSAO adding depth and contact shadows to terrain geometry
Screen-space ambient occlusion adding contact shadows and depth cues to voxel terrain geometry

Bloom and tile atlas

The bloom system uses a two-pass tile atlas driven by compute-generated mipmaps, split across composite4 (blur atlas generation) and composite5 (compositing). Instead of running multiple downsampling passes, the compute mipmap generator produces pre-filtered levels, which are packed across seven LODs (LOD 2 through 8, representing 1/4 to 1/256 resolution) into a single atlas in colortex3.

In composite4, each tile samples its corresponding mip level from colortex0 and applies a 7x7 Gaussian blur with Pascal’s triangle row-6 weights (1, 6, 15, 20, 15, 6, 1) normalized by 4096. Because each mip texel represents exp2(lod) screen pixels, single-texel offsets at mip scale produce a wide filter footprint. The blurred output is gamma-encoded as pow(x/128, 0.25) to fit HDR values into the RGBA8 format of colortex3.

In composite5, GetBloomTile() unpacks each tile, decodes the values (x^4 * 128), and DoBloom() averages all seven levels with equal weight (1/7). A darkness boost factor scales up bloom intensity in dark environments, and tile layouts normalize to 1920x1080 to prevent tile overlap at smaller display resolutions.

graph LR
    CT0["colortex0 HDR"] --> MIP["Compute Mipmap Generator"]
    MIP --> C4["composite4: 7x7 Gaussian per LOD"]
    C4 --> ATLAS["colortex3 Tile Atlas (7 LODs)"]
    ATLAS --> C5["composite5: Read + Decode + Average"]
    C5 --> BLOOM["Bloom + Scene Blend"]
Bloom effect with mipmap-based tile atlas
Bloom effect showing soft glow around bright light sources via 7-LOD tile atlas and Gaussian blur

Tonemapping and color grading

The final color stage in composite5 applies Lottes 2016 tonemapping followed by BSL-style color grading. The Lottes operator provides exposure, contrast, and highlight roll-off adjustments through polynomial coefficients (a, b, c, d).

The tonemapping pipeline runs through seven adjustments: exposure scaling, Lottes polynomial evaluation, linear-to-sRGB gamma conversion (IEC 61966-2-1), dark lift (a smoothstep blend that lightens deep shadows for contrast), white-point compression for bright highlights, and shadow desaturation to prevent color clipping in dark pixels. Settings are exposed through settings.hlsl.

Following tonemapping, color grading adjusts saturation and vibrance. Vibrance calculates a per-pixel saturation metric from channel spreads to boost muted tones while leaving saturated colors stable, while global saturation shifts colors relative to perceptual gray. Default settings apply a 7% saturation lift.

Sky glare

The sky pass (gbuffers_skybasic) adds an atmospheric halo around the sun and moon. Glare intensity uses an inverse Fresnel formula based on view-to-light alignment, with an adaptive scatter exponent that concentrates the core while preserving outer falloff. Rain broadens and dims the glare while shifting its tint toward neutral grey to simulate cloud scattering.

Sky glare halo around the sun at golden hour
Atmospheric sky glare around the sun, with Fresnel-driven halo intensity and warm color transition

Glare colors shift from sunset gold to midday white for the sun, and cool blue-grey for the moon, with a 7x brightness scale applied underwater. Parameters such as SUN_GLARE_VISFACTOR (width) and SUN_GLARE_STRENGTH (intensity) can be configured per quality profile in settings.hlsl.

Design decisions

Engine and application separation

The engine manages GPU resources, bindless descriptor heaps, and stateless drawing APIs, while the application layer in Game::RenderWorld() defines pass order and execution flow. This keeps the engine core independent of any specific pipeline architecture, allowing it to drive forward, deferred, or hybrid pipelines without changes.

Data-driven shader configuration

Render target formats, blend modes, draw buffers, material mappings, quality presets, and custom texture bindings are declared in shaders.properties, block.properties, comment directives, and HLSL constants. The engine’s parsers (CommentDirectiveParser, ConstDirectiveParser, ShaderProperties) translate these into runtime state during bundle loading, creating PSOs and configuring render targets without modifying C++ code.

Resource binding and throughput

GraphicsRootBinder tracks binding validity to skip redundant D3D12 root signature updates, and PerPass scopes prevent constant buffer re-uploads across composite stages. Chunk batching with frustum culling aggregates geometry into 4x4 regions to cut draw calls before submission. Multi-frame in-flight pipelining overlaps CPU recording with GPU execution, and compute-shader mipmap generation produces downsampled tiers for bloom in place of separate blur passes.

Fallback chains and runtime reloading

The fallback hierarchy (user sub-bundle, user bundle, engine bundle) ensures every pass resolves to a valid shader program, even if an experimental bundle omits specific passes. MulticastDelegate lifecycle events and frame-boundary bundle swapping allow authors to reload and test shaders while the engine is running.