The paradigm of corporate B2B events has fundamentally shifted. Passive webinar formats no longer meet the engagement demands of a discerning enterprise audience. To deliver impactful keynotes, product launches, and all-hands meetings in a hybrid world, forward-thinking organizations are adopting broadcast-level technologies. Among the most transformative of these is 3D virtual production, a workflow that moves beyond simple green screens and into the realm of photorealistic, real-time rendered environments. Executing this at an enterprise scale requires more than creative software; it demands a robust, meticulously engineered media hub infrastructure. This article provides a technical deep-dive into the end-to-end workflow of a 3D virtual production, from network architecture to final pixel delivery, designed for the AV professionals, IT directors, and production managers tasked with its implementation.
Core Infrastructure: The Media Hub Foundation for Virtual Production
A successful virtual production is built upon a high-performance technical foundation. The immense data throughput and low-latency requirements of real-time rendering and uncompressed video transport necessitate an infrastructure far beyond a standard enterprise IT environment. This core is the central nervous system of the entire operation, where compromises in network, compute, or synchronization result in catastrophic failure of the live production.
Network Fabric: Precision Timing and High-Bandwidth Architecture
The network is arguably the most critical component. We are not merely transmitting data; we are transporting time-sensitive, uncompressed video and tracking information that must remain in perfect lockstep. A standard 1GbE network is wholly insufficient. The baseline for professional virtual production is a 10GbE network, with 40GbE or 100GbE in the core or for specific high-bandwidth links. We architect these networks using a spine-and-leaf topology, which provides predictable, low-latency paths between any two points and avoids the bottlenecks of traditional three-tier architectures. To ensure all devices share a common time reference, which is essential for video synchronization, the network must support Precision Time Protocol (PTP) as defined by SMPTE ST 2059. This protocol provides sub-microsecond accuracy, allowing cameras, render engines, and switchers to be genlocked over IP without requiring separate analog black burst cabling.
Compute and GPU Clustering for Real-Time Rendering
The heart of the 3D environment is the real-time render engine, most commonly Epic Games’ Unreal Engine. Rendering a complex, photorealistic scene at a target of 59.94 frames per second requires immense computational horsepower. This is achieved through dedicated servers equipped with high-end NVIDIA GPUs, such as the RTX A6000 or RTX 6000 Ada Generation, which are designed for professional visualization and feature support for crucial technologies like NVIDIA GPUDirect for Video. For large-scale LED wall productions, multiple render nodes are often clustered using Unreal Engine’s nDisplay technology. This allows different sections of the virtual world to be rendered by different machines in parallel, creating a seamless, ultra-high-resolution canvas. These servers require fast, local storage, typically NVMe RAID arrays, to load high-resolution textures and complex 3D assets without introducing latency during a live session.
Ingest, I/O, and Synchronization Hardware
Connecting the physical world to the virtual one requires specialized input/output (I/O) hardware. Professional video capture cards from manufacturers like AJA or Blackmagic Design are used to ingest camera feeds into the render engine servers. These feeds are typically 12G-SDI for 4K UHD signals. Crucially, every video source in the production, including the physical cameras and the output of the render engines, must be synchronized to a master clock. A dedicated master sync generator provides a reference signal (either tri-level sync for HD/UHD or PTP over the network) to all devices, preventing frame tearing and motion artifacts that would otherwise destroy the illusion of the virtual set.

The Real-Time Rendering Pipeline: From Unreal Engine to Live Output
With the infrastructure in place, the focus shifts to the software and data pipeline that brings the virtual world to life. This is a sequence of precise, interconnected processes where data from the real world (camera position, lens settings) is used to generate a corresponding virtual view, which is then composited with the live talent.
Camera Tracking and Lens Calibration
For the virtual environment to react realistically to physical camera movements, its position and orientation must be tracked in real-time with millimeter accuracy. Systems like Mo-Sys StarTracker (optical) or stYpe use infrared sensors to track markers on the camera, while other solutions rely on protocols like FreeD to send pan, tilt, zoom, and focus data directly from the camera head. This tracking data is sent over the network to the Unreal Engine server. Just as critical is lens calibration. Every physical lens has unique distortion and field-of-view (FOV) characteristics. Before production, each lens is carefully calibrated to create a digital profile. The render engine uses this profile to ensure the virtual camera’s properties perfectly match the physical one, eliminating any visual disconnect between the real foreground and the virtual background.
Real-Time Compositing and Green Screen Keying
Once the render engine receives the tracking data, it generates the correct virtual background for the camera’s current perspective. This rendered image must then be combined with the live-action talent. If using a green screen, the SDI feed from the camera is ingested into a dedicated hardware or software keyer. The keyer removes the green background, and the resulting foreground plate of the talent is then composited over the virtual background inside Unreal Engine or in a downstream production switcher. For more advanced LED volume stages, the rendered environment is displayed directly on the LED wall behind the talent. The camera captures both the talent and the background in a single shot, a technique known as in-camera VFX.
Signal Integrity and Transport: Ingest, Routing, and Distribution
Managing the flow of numerous high-bandwidth signals is a core challenge. A typical production involves multiple camera feeds, camera tracking data streams, rendered outputs from several GPU servers, audio channels, and final program feeds. This requires a robust routing and switching architecture to maintain signal integrity and production flexibility.
Uncompressed vs. Mezzanine-Compressed Workflows
The choice of video transport protocol has significant architectural implications. For the highest possible quality and near-zero latency, an uncompressed workflow using SMPTE ST 2110 is the broadcast industry standard. This standard transports video, audio, and ancillary data as separate streams over an IP network. However, it requires a highly managed 40/100GbE network and ST 2110-specific hardware. A more accessible yet still high-quality alternative is using a mezzanine compression codec like NDI (Network Device Interface). NDI offers visually lossless quality at much lower bandwidths, allowing it to run over a standard 10GbE network. While NDI introduces a few frames of latency, its ease of use and software-based approach make it a powerful choice for many corporate media hubs.
The Role of the Master Production Switcher
The output from the Unreal Engine render nodes, which represents the final composited virtual scene, is treated as a video source just like any physical camera. These sources are fed into a master production switcher (vision mixer), such as a Ross Video Carbonite or Grass Valley Kula. The technical director uses this switcher to cut between different virtual cameras, physical cameras in the studio, and insert other production elements like presentation graphics or pre-recorded video packages. The switcher outputs the final program feed, which is the polished content intended for the audience. Multiview monitors connected to the switcher allow the production crew to see all sources, preview, and program outputs simultaneously.

Audio Integration and Synchronization
Audio is an equal partner to video. Audio from microphones on set is typically embedded into the camera’s SDI signal or transported separately over the network using protocols like Dante. It is crucial to manage the audio path to account for the video processing latency introduced by the render engine. An audio delay is applied to the microphone signals within the audio mixer to ensure perfect lip-sync in the final program output. For hybrid events with remote presenters, a mix-minus feed is generated for each remote participant. This feed contains the full program audio minus their own voice, preventing echo and feedback.
Hybrid Event Integration: Synchronizing Physical and Virtual Audiences
The final stage of the workflow is delivering the pristine program feed from the media hub to both physical and virtual audiences without compromising quality. This involves robust encoding, strategic integration with enterprise collaboration platforms, and broadcast-level redundancy.
Encoding for Contribution and Distribution
The program output from the production switcher (typically as a 12G-SDI signal) is fed into a professional hardware encoder. For sending the feed to a cloud video platform or another production facility, the Secure Reliable Transport (SRT) protocol is the preferred method. SRT provides low-latency, high-quality video transport over unpredictable public networks like the internet. For direct distribution to the audience, the stream is encoded using efficient codecs like H.264 or H.265 (HEVC) into an adaptive bitrate (ABR) ladder. This ABR stream is then sent to a Content Delivery Network (CDN), which ensures a smooth viewing experience for viewers on different devices and network conditions.
Integrating with Enterprise Platforms
A common requirement for B2B events is to stream into platforms like Microsoft Teams, Zoom Events, or Webex. Directly feeding a high-bitrate SRT stream is often not possible. The professional workflow involves using a high-quality hardware or software output from the production system that can be recognized as a virtual webcam source by these platforms. Alternatively, platforms like Teams can use NDI inputs, allowing a direct feed from the production network. It is critical to manage the encoding and transport to these platforms carefully to maintain as much quality as possible, providing a far more professional result than a simple screen share.
Ultimately, navigating the workflow of a 3D virtual production is a complex but achievable engineering challenge. It requires a holistic approach that balances creative vision with a deep understanding of network engineering, video transport protocols, and real-time graphics processing. By building on a foundation of broadcast-grade infrastructure and meticulous planning, organizations can leverage this technology to create truly immersive and impactful B2B events that captivate both in-person and remote audiences. At Spring Forest Studio, our technical team specializes in designing and implementing these sophisticated workflows, ensuring flawless execution for mission-critical corporate communications.

Jeremy Lee is a seasoned digital marketing director and strategist with over two decades of experience in the industry. As the founder of Sotavento Medios, I manage a diverse portfolio of over 50 businesses, helping brands grow through advanced search strategies and digital innovation. My work focuses on bridging the gap between traditional search engine optimisation and the evolving world of AI-driven answer engines.
get in touch