Profiling
Engine-owned documentation for reusable Tracy instrumentation, capture boundaries, and comparable performance measurements. Workload scenes, launch orchestration, acceptance budgets, and retained reports belong to the embedding game project.
Purpose
Use this guide when a client, server, tool, or baker is slow and a timing capture is more useful than a debugger trace. The Engine supplies Tracy instrumentation and dedicated build configurations. A game project supplies the process topology and a deterministic workload.
The central rule is simple: instrument only the process being measured. A profiled client should connect to a regular server; a profiled server should be driven by a regular client. This keeps one Tracy endpoint, one process timeline, and one attribution boundary per capture.
The Regular counterpart participates through the ordinary game client/server connection. It neither opens nor owns the Tracy endpoint and never serves as a placeholder holder for the Tracy port.
Capture decision
- Select
Profiling_OnDemandfor a bounded manual capture orProfiling_Totalwhen startup must be included. On a single-config generator, choose the profiling configuration at configure time rather than expecting a later--configswitch to replace it. - Run one Profiled process and a Regular counterpart so only the measured process owns the default Tracy port and timeline. The Regular counterpart neither opens nor owns a Tracy endpoint. For a client capture, start the Regular server and wait for readiness, reject a stale instrumented process immediately before launching the Profiled client, then launch that client. For a server capture, reject a stale instrumented process immediately before launching the Profiled server, wait for readiness, and only then drive it with the Regular client.
- Fix revisions, configuration, workload, input, map/data state, duration, and warm-up condition. Record one capture at a time and repeat the same case at least three times before comparing a representative result.
The role-specific startup order is exact:
| Capture | Ordered startup |
|---|---|
| Client | Start the Regular server and wait for readiness; reject a stale process on the default Tracy port; launch the Profiled client; then start tracy-capture. |
| Server | Reject a stale process on the default Tracy port; launch the Profiled server and wait for readiness; start the Regular client workload driver; then start tracy-capture. |
The stale-port check occurs immediately before the Profiled process in both routes. Never move it earlier than the Regular-server readiness gate in the client route or after the Profiled process merely because the capture tool itself starts later.
Source paths inspected
BuildTools/cmake/stages/Init.cmakeBuildTools/cmake/stages/ThirdParty.cmakeSource/Essentials/BasicCore.hSource/Essentials/StackTrace.hSource/Essentials/Logging.cppSource/Essentials/MemorySystem.cppSource/Essentials/Threading.hSource/Frontend/ApplicationInit.cppSource/Frontend/Application.cppSource/Frontend/ApplicationHeadless.cppSource/Client/Client.cppSource/Server/Server.cppSource/Scripting/AngelScript/AngelScriptContext.cppSource/Applications/BakerLib.cppThirdParty/tracy/CMakeLists.txtThirdParty/tracy/NEWSThirdParty/tracy/public/common/TracyVersion.hppExamples/MinimalMultiplayer/CMakeLists.txtExamples/MinimalMultiplayer/CMakePresets.jsonExamples/MinimalMultiplayer/FOnlineMinimalMultiplayer.fomainExamples/MinimalMultiplayer/README.mdExamples/MinimalMultiplayer/run_tutorial_smoke.py- upstream Tracy v0.13.1
capture.cppandcsvexport.cpp
What the Engine instruments
Profiling configurations compile TracyClient into the Engine essentials
layer and define FO_TRACE_ENABLED. Other configurations compile zones away.
FO_TRACE_CATEGORIES selects which categories are built into a Tracy binary:
empty (the default) enables all zone categories but neither opt-in signal;
Render,Gui keeps only those zones, +Memory adds allocation events, and
-Script removes script zones. Unknown names fail configuration. The selected
0/1 values are generated into TraceCategories.gen.h.
| Signal | Current Engine behavior |
|---|---|
| Process identity | ApplicationInit sends FO_NICE_NAME to Tracy. |
| Native CPU zones | FO_TRACE_ZONE(Category) and FO_TRACE_ZONE_NAMED(Category, name) emit colored zones only for selected categories. They are not native stack traces. |
| Script CPU zones | The Script category gates generated bindings, AngelScript calls, and Mono method zones. AngelScript resumes its script zone stack after suspension. |
| Frames | Visible and headless Application::EndFrame() paths emit FrameMark. |
| Client plot | Client FPS is emitted from ClientEngine::MainLoop(). |
| Server plot | Server jobs per second is emitted from the server statistics job. |
| Log messages | Engine log records become Tracy messages only with opt-in Log. |
| Threads | Engine thread names are visible in Tracy and in tagged log lines. |
| Memory | Engine allocator events require opt-in Memory; ordinary C-library allocations remain outside this accounting. |
Zone categories declared in Source/Essentials/BasicCore.h are App,
Engine, Entity, Map, Script, Network, Database, Threading,
Render, Model, Particles, Gui, Audio, FileSystem, Core, Baking,
and Editor. Frame marks, plots, and thread names remain enabled in every
Tracy build. FO_TRACE_CATEGORY_ENABLED(Category) allows category-dependent
hooks; a non-Tracy build still validates the category name.
The current first-party integration does not add renderer GPU zones or Tracy lock wrappers. A CPU capture therefore must not be reported as GPU timing or lock-contention proof. Use renderer-specific tools or deliberately add scoped instrumentation when that evidence is required.
Plain C-library allocations made outside the Engine allocation wrappers are not automatically attributed to the Engine heap. See Essentials before interpreting an allocation capture as complete process memory accounting.
Build configurations
BuildTools/cmake/stages/Init.cmake adds four configurations:
| Configuration | Base | Tracy mode | Use |
|---|---|---|---|
Profiling_OnDemand |
RelWithDebInfo |
on demand | Normal captures after startup and warm-up. |
Profiling_Total |
RelWithDebInfo |
continuous | Startup, initialization, and first-frame captures. |
Debug_Profiling_OnDemand |
Debug |
on demand | Diagnose instrumentation or debug-only behavior, not a performance baseline. |
Debug_Profiling_Total |
Debug |
continuous | Diagnose debug startup with full tracing. |
On-demand mode defines TRACY_ON_DEMAND, so tracing starts only after a
profiler connection. It is the default choice for repeatable steady-state
measurements. Total mode records from process startup and is appropriate only
when startup is part of the question. Restart a total-mode process before a
new capture.
For a multi-config generator, select the profile with cmake --build
--config. For a single-config generator, set CMAKE_BUILD_TYPE to the
profiling configuration at configure time. Passing --config to a
single-config build does not turn an already configured release binary into a
Tracy binary.
Prepare matching Tracy tools
The Engine currently vendors Tracy 0.13.1. The checkout is intentionally
pruned to the client library under ThirdParty/tracy; it does not contain the
upstream capture, csvexport, or GUI profiler sources. The capture protocol
must match, so use the
Tracy v0.13.1 release
or build tools from that exact tag.
The documented capture and export flags are pinned to the same tag’s
capture.cpp
and csvexport.cpp,
not inferred from a newer local installation.
The following creates headless capture and CSV-export tools outside authored source:
cmake -E make_directory Workspace/Tracy
git clone --depth 1 --branch v0.13.1 https://github.com/wolfpld/tracy.git Workspace/Tracy/src
cmake -S Workspace/Tracy/src/capture -B Workspace/Tracy/capture -DCMAKE_BUILD_TYPE=Release -DNO_FILESELECTOR=ON
cmake --build Workspace/Tracy/capture --config Release
cmake -S Workspace/Tracy/src/csvexport -B Workspace/Tracy/csvexport -DCMAKE_BUILD_TYPE=Release -DNO_FILESELECTOR=ON
cmake --build Workspace/Tracy/csvexport --config Release
NO_FILESELECTOR=ON avoids pulling GUI file-dialog dependencies into the two
headless command-line tools. The output directory differs between
single-config and multi-config generators; locate tracy-capture and
tracy-csvexport after the build and put them on PATH. Use the matching
release’s GUI profiler to inspect saved .tracy files.
Do not silently use a different tool release. Tracy rejects an incompatible handshake, and a tool that can open an older file is not proof that its live capture protocol matches the instrumented process.
Choose one measurement boundary
| Question | Profiled process | Regular counterpart |
|---|---|---|
| Rendering, visibility, UI, input, client scripts, or client networking | desktop client | headless or desktop server |
| Authority, AI, simulation, persistence, server scripts, or server networking | headless or desktop server | standalone client workload driver |
| Resource-baking throughput | baker application | no runtime counterpart |
| Startup and first frame | affected process in Profiling_Total |
only the minimum dependencies needed to reach startup |
Do not profile an embedded server and client in the same process when the question requires client/server attribution. Do not build both standalone processes with Tracy on the default port and then guess which one accepted the capture connection.
Build and bake before the measurement window. An unexpected in-process rebake, shader warm-up, cache population, or updater operation is part of the capture only when that operation is the stated workload.
Build a profiled sample
The engine-owned minimal multiplayer project gives the commands concrete target names without depending on a private game. On Windows:
Set-Location Examples\MinimalMultiplayer
cmake --preset windows
cmake --build Build\windows --config RelWithDebInfo --target BakeResources FOMM_ServerHeadless
cmake --build Build\windows --config Profiling_OnDemand --target FOMM_Client
On Linux:
cd Examples/MinimalMultiplayer
cmake --preset linux
cmake --build Build/linux --config RelWithDebInfo --target BakeResources FOMM_ServerHeadless
cmake --build Build/linux --config Profiling_OnDemand --target FOMM_Client
This prepares a profiled client and a regular server. To profile the server, reverse the configurations:
cmake --build Build/linux --config RelWithDebInfo --target BakeResources FOMM_Client
cmake --build Build/linux --config Profiling_OnDemand --target FOMM_ServerHeadless
Project target names are derived from FO_DEV_NAME; FOMM_* names are valid
only for this engine-owned sample.
Capture a client
Use separate terminals and make Examples/MinimalMultiplayer/Build/windows
the runtime working directory, matching the generated target and smoke-test
contract. Start the regular server first:
Set-Location Examples\MinimalMultiplayer\Build\windows
.\Binaries\Server-Windows-win64\FOMM_ServerHeadless.exe -ApplyConfig ..\..\FOnlineMinimalMultiplayer.fomain
Start the profiled client:
Set-Location Examples\MinimalMultiplayer\Build\windows
.\Binaries\Client-Windows-win64-Profiling_OnDemand\FOMM_Client.exe -ApplyConfig ..\..\FOnlineMinimalMultiplayer.fomain
From the example project root, reach a stable workload and capture a fixed window:
New-Item -ItemType Directory -Force -Path Workspace\Profiling | Out-Null
tracy-capture -o Workspace\Profiling\client.tracy -f -s 30 -m 80 -p 8086
For Tracy 0.13.1, -f permits replacing the named output, -s 30 captures for
30 seconds, -m 80 caps capture memory at 80 percent of physical memory, and
-p 8086 selects the default Tracy port. Keep capture stdout with the result;
it records the frame count, time span, zone count, and trace size.
The Linux binary layout follows the same rule:
Binaries/Server-Linux-x64/ for the regular server and
Binaries/Client-Linux-x64-Profiling_OnDemand/ for the profiled client.
Capture a server
Build FOMM_ServerHeadless as Profiling_OnDemand and FOMM_Client as
RelWithDebInfo. Start the profiled server, connect the regular client, wait
until the selected workload is stable, and run the same tracy-capture
command.
The profiled Windows server is emitted below
Binaries/Server-Windows-win64-Profiling_OnDemand/; the regular client remains
below Binaries/Client-Windows-win64/.
For a startup capture, build the target as Profiling_Total, start
tracy-capture before the process, and retain the complete startup log. Do not
compare a total-mode startup trace with an on-demand steady-state trace.
Design a reproducible workload
A performance result needs a workload contract, not only a .tracy file.
Record:
- exact Engine and game revisions, dirty-worktree state, target, and profile;
- compiler, host CPU, operating system, renderer, driver, resolution, and frame cap;
- map/scene or server workload ID, content revision, seed, actor/client count, and account/database setup;
- warm-up condition, capture duration, frame or server-job count, and all input automation;
- other processes, host CPU load, thermal/power mode, and any discarded attempt;
- raw capture, capture stdout, runtime logs, exported tables, and the interpretation.
Prefer a project-owned scripted scene or workload driver. It should start from a known state, suppress unrelated human input, announce readiness in a log, run the same actions for every attempt, and stop cleanly. Keep visual debug overlays, admin panels, verbose diagnostics, and unrelated telemetry out of the profile unless they are the subject.
Run one capture at a time. Check that the Tracy port is free and that no stale instrumented process can accept the connection. Wait for an idle host, reject captures with unexpected frame/job counts, and retain at least three comparable attempts when making an optimization claim. Report the selected statistic and selection rule instead of keeping only the fastest trace.
Analyze a capture
Start with the timeline and call tree:
- Confirm program name, process, build configuration, capture duration, and workload marker.
- Confirm frame or server-job counts are comparable to the baseline.
- Find long frames, scheduling gaps, and dominant threads before sorting functions.
- Compare inclusive time with self time. A large parent zone can merely own expensive children.
- Inspect AngelScript zones in the same timeline as their native callers.
- Correlate log messages, FPS/job plots, allocations, and source locations with the time window.
- Change one variable, repeat the same workload, and preserve both raw captures.
The matching CSV exporter can produce a source-linked self-time table:
tracy-csvexport -e --truncated_mean=95 Workspace/Profiling/client.tracy > Workspace/Profiling/hotspots-self.csv
-e selects self time. The 95-percent truncated mean reduces the influence of
the slowest tail while retaining a separate percentile column; it does not
replace inspection of long frames or tail latency. Use the same export options
for baseline and candidate captures.
Client captures expose the Client FPS plot. Server captures expose Server
jobs per second; the current event-driven server no longer publishes the
removed loop-time/loops-per-second metrics. See
Server Runtime for the current
server statistics boundary.
Managed script zones
With the Script category enabled, the Managed C# backend uses the Mono profiler to make each
instrumented game-script method a Tracy zone carrying its source file and
declaration line. Native Engine zones nest below it, so a log line such as
Script execution overrun: GameScripts.Audio.OnLoop identifies the entry to
find, while the capture identifies the work underneath it:
GameScripts.Audio.OnLoop
GameScripts.Audio.TryPlay
AudioManager::PlayMusic
AudioManager::Load
FileSystem::ReadFile
Compilation is a separate JIT zone over the method being compiled. It is not
restricted to game assemblies because a handler’s first invocation pays for
every runtime method it reaches. Call zones, by contrast, cover only registered
game assemblies; generated wrappers and runtime-library frames are deliberately
absent. An inlined method may therefore contribute time without a named zone.
Choose the capture mode for the question. Profiling_OnDemand starts recording
after startup and normally contains neither OnStart* work nor most first-use
JIT. Use Profiling_Total when startup is the workload. Method instrumentation
also prevents some inlining and adds an entry/leave hook, so tiny accessors are
the most distorted zones. Call counts remain exact; treat absolute managed-zone
times as an upper bound and compare like-for-like captures rather than comparing
them directly with RelWithDebInfo.
The profiler-hook lifecycle, image filter, tail-call exclusion, and shared method metadata cache are documented in Managed C# Scripting.
Placing zones
Use existing zones before adding new ones. Put FO_TRACE_ZONE(Category) as
the first statement in selected nontrivial .cpp functions: frame/tick
orchestration, variable work, blocking operations, or rare expensive paths.
Use FO_TRACE_ZONE_NAMED(Category, name) for generated bindings, event
subscriptions, or a meaningful job lambda. Generated script exports get a
binding zone rather than another direct zone. Split cheap common work from an
expensive conditional region when appropriate (for example, entity lock wait
or sprite flush). Do not instrument accessors, trivial forwarders, hot
per-element helpers, constructors/destructors, profiler internals,
non-returning loops, headers, lambdas, or tests indiscriminately; the
EntityEventWrapper::FireSubscribed header zone is a deliberate exception.
Guard direct Tracy macros with #if FO_TRACE_ENABLED, and use the Memory
or Log opt-in only when that signal is needed.
Instrumentation must:
- compile out in non-profiling configurations;
- use stable, low-cardinality names;
- avoid secrets, account data, chat text, or unbounded content IDs;
- avoid changing ownership, locking, allocation, or scheduling semantics;
- include a focused validation route and an update to this guide when it changes the public interpretation of a capture.
Do not add a zone solely to make a report look more detailed. Each zone should answer a named performance question and have an owner who can interpret it.
Common failure modes
| Symptom | Check |
|---|---|
| Capture tool cannot connect | Confirm the process was built in one of the four profiling configurations, the port is correct, and a firewall or stale process is not owning it. |
| Protocol mismatch | Use tools from the exact version in TracyVersion.hpp. |
| Trace starts too late | Use Profiling_Total only when startup is the intended workload; otherwise add an explicit readiness marker and warm-up. |
| Near-empty trace | Confirm the capture attached to the intended process and that the workload continued for the full window. |
| Client and server zones are mixed or ambiguous | Profile one standalone process and rebuild the counterpart as RelWithDebInfo. |
| Profile is dominated by baking or cache creation | Pre-bake and warm up, or rename the workload as a startup/baking measurement. |
| Debug build looks much slower | Use Profiling_OnDemand or Profiling_Total for performance comparisons; debug profiles are diagnostic. |
| Headless client result is used as renderer evidence | A null/headless renderer does not validate visible rendering performance. |
| Software rendering dominates Linux capture | Record the renderer and driver; do not compare software and hardware renderer captures. |
| Allocation totals appear incomplete | Check for third-party/plain C allocations outside the Engine allocator boundary. |
Allocator occupancy without a Tracy capture
Debug and Tracy builds with rpmalloc expose occupancy automatically. For a regular
build, opt in with FO_MEMORY_DIAGNOSTICS=ON; the default remains OFF. Native
memory::get_allocator_statistics() and the shared AngelScript/Managed C#
Game.GetAllocatorStatistics() distinguish global module pages from calling-thread
size classes. An empty script dictionary means unavailable, not an empty heap.
This does not enable Tracy capture or measure whole-process/GPU fragmentation.
Interpret capacities and the non-atomic sampling boundary using
Essentials.
Project-owned automation
An embedding project should automate the repeatable parts without changing the Engine contract. A production runner should:
- choose a declared workload ID and exact client/server measurement side;
- build only the measured process with Tracy;
- pre-bake and stage binaries/resources;
- serialize access to runtime logs and the Tracy port;
- wait for process and workload readiness;
- enforce quiet-host and minimum-work checks;
- run
tracy-capture, export a stable table, and copy all logs; - retain provenance and reject noisy or incomplete attempts.
Scene IDs, config names, executable names, expected FPS/job thresholds, database fixtures, renderer policy, and report destinations remain project-owned. They must not be copied into this Engine guide as defaults.
Validation workflow
After changing profiling configurations, instrumentation, or this guide, run:
python BuildTools/tests/test_docs_profiling.py
python BuildTools/docs_snippets.py --check --external
python BuildTools/docs_validate.py
Then build one on-demand target and verify a short capture contains:
- the expected program name and process;
- native and AngelScript zones;
- frame marks;
- the applicable client/server plot;
- log messages and named threads;
- a nonzero frame/job count and a readable saved trace.
Memory changes also require a focused allocation capture. Script-context changes require a capture that executes, suspends, resumes, and completes an AngelScript call chain. Frontend changes require both visible and headless frame-mark paths when both remain supported.
Maintenance
Update this page in the same change when:
- profiling configuration names or their base configurations change;
FO_TRACE_ENABLED,FO_TRACE_CATEGORIES,TRACY_ENABLE, orTRACY_ON_DEMANDwiring changes;- the vendored Tracy version or pruned payload changes;
- frame marks, program/thread naming, logs, plots, allocation tracking, or AngelScript zones change;
- a first-party GPU/lock instrumentation boundary is added;
- the engine-owned sample target/config/output layout changes.
When updating Tracy, verify the client and matching upstream capture/export tools together. Re-run one client and one server capture before changing the documented version or protocol claim.
See also
- Testing for test-boundary selection, sanitizers, and coverage.
- Native, AngelScript, and Managed Debugging for native and script diagnostics/debugger workflows.
- Build Workflow for embedding-project build ownership.
- Frontend and Rendering for renderer/runtime boundaries.
- Server Runtime for current job-throughput diagnostics.
- ThirdParty Maintenance for vendored version and pruning updates.