A step-by-step guide has been published, detailing how to set up a programmable pipeline for generating both video and audio content using the MiniMax-H3 multimodal model. The pipeline utilizes ComfyUI as a backend to automate tasks such as hardware profiling and model weight downloading. The guide covers the process of constructing a dynamic graph and decoding joint video-audio content. This implementation is aimed at developers looking to leverage multimodal generation capabilities in their applications. The development of such pipelines has the potential to enhance the creation of immersive and interactive content.