AI/ML Engineer — GenAI (Image, Video, Audio-Performance & 3D), Mid-Level
Experience: 4-6 Years
Department:AI/ML — Generative Media
Location: Hderabad
About the Role
As a Mid-Level AI/ML Engineer, you will build and fine-tune current-generation image, video, audio-performance, and 3D generative models that feed directly into our VFX and game content pipelines. You will work closely with senior engineers, technical artists, and pipeline TDs to take models from experimentation to production-ready assets and tools. You are expected to be comfortable downloading open-weight checkpoints, understanding their architecture, and modifying them — not just wrapping them in inference scripts.
Key Responsibilities
- Train and fine-tune current-generation image and video diffusion transformers across the FLUX (.1/.2 dev/schnell/klein), Qwen-Image, Z-Image (Base/Turbo/De-Distilled), Krea 2 (RAW/Turbo), Boogu (Base/Turbo/Edit), and Wan (2.x) families, plus unified multimodal generation/editing/understanding models (JoyAI-Image, BAGEL, OmniGen2-class architectures).
- Build and maintain fine-tuning and conditioning pipelines using modern parameter-efficient methods (LoRA, DoRA, LyCORIS / LoHa / LoKr) and current-generation training backends that support layer-specific training, block-wise optimization, fp8 scaling, and low-VRAM full fine-tuning for style, character, and asset consistency.
- Build audio-driven performance capability — lip-sync and talking-head/digital-human generation using audio-conditioned diffusion approaches in the OmniSync, Hallo3, MuseTalk, LatentSync, and Sonic lineage, including mask-free DiT approaches and native audio-video diffusion (LTX-2.3).
- Prepare, clean, and curate training datasets, including captioning/labeling pipelines for images, video clips, and paired audio-visual data.
- Implement and tune training loops (flow-matching/rectified-flow objectives, schedulers, EMA, mixed precision) and monitor convergence and quality metrics.
- Support multi-GPU training runs and help optimize training/inference throughput on the studio GPU cluster.
- Optimize inference pipelines for latency and memory (FP8/INT8/INT4 quantization, attention optimization, ONNX/TensorRT export, torch.compile) for integration into artist-facing tools.
- Train and fine-tune current-generation 3D generative models (TRELLIS 2, Hunyuan3D, Tripo-class image/text-to-3D) for game-ready assets and environment props, including mesh, texture, and PBR material generation.
- Support experimentation with neural scene representations (NeRF, 3D Gaussian Splatting) for environment reconstruction and generation workflows.
- Create and maintain quantized model artifacts (GGUF, FP8, INT8) for deployment on varied hardware tiers, including converting and quantizing custom fine-tuned checkpoints for diffusers-based and ComfyUI-GGUF pipelines.
- Track experiments, document findings, and contribute to internal best practices for GenAI model development.
- Collaborate with pipeline/tools engineers to integrate models into the studio creative pipelines.
Required Skills & Experience (Mandatory)
The following are mandatory. We will verify hands-on experience during technical interviews.
- Python, PyTorch (primary), and familiarity with JAX/TensorFlow.
- Hands-on, open-source-first experience: comfortable downloading, running, fine-tuning, and modifying open-weight models directly (Hugging Face Hub, GitHub repos) — not solely prompting or orchestrating hosted commercial APIs (e.g. OpenAI, Anthropic, Runway).
- Hugging Face Diffusers, Transformers, and PEFT libraries — you should know how to patch a pipeline, swap a scheduler, and debug a forward pass.
- Current image generation architectures: Diffusion Transformer / rectified-flow image models — FLUX.1/FLUX.2 (dev/schnell/klein), Qwen-Image, Z-Image, Krea 2, Boogu (Base/Turbo/Edit) — and working awareness of unified multimodal generation/editing/understanding models (JoyAI-Image, BAGEL, OmniGen2-class).
- Current video generation architectures: Mixture-of-Experts video DiT (Wan-class) and audio-video native DiT (LTX-2-class), plus the broader DiT lineage (HunyuanVideo, CogVideoX).
- Training objectives: flow matching and rectified flow as the primary paradigm, alongside classical DDPM/DDIM/EDM schedulers for context.
- Fine-tuning and conditioning: LoRA, DoRA, LyCORIS (LoHa, LoKr), and modern conditioning approaches for DiT models (in-context learning / visual prompting, VLM hidden-state conditioning, style-reference systems, AdaLN modulation, and native MM-DiT joint attention — NOT legacy IP-Adapter or ControlNet as primary methods).
- Working exposure to audio-driven generation: at least conceptual or hands-on familiarity with one audio-conditioned lip-sync/talking-head system (OmniSync, Hallo-family, MuseTalk, LatentSync, Sonic, or comparable).
- Current 3D generative model families: image/text-to-3D systems such as TRELLIS 2, Hunyuan3D (2.x/3.x), or Tripo, producing textured, PBR-ready meshes; plus NeRF and 3D Gaussian Splatting fundamentals.
- Mixed-precision training (FP16/BF16/FP8), gradient checkpointing, EMA, and basic multi-GPU training (DDP/FSDP) on single or multi-node clusters.
- Model export & optimization basics: ONNX, TensorRT, quantization (FP8/INT8/INT4), and optimized attention kernels (FlashAttention, SageAttention, or equivalent).
- Experiment tracking (Weights & Biases / MLflow) and dataset curation/captioning pipelines.
- Version control (Git) and Linux/CUDA environment fundamentals.
- Solid understanding of diffusion/flow-matching theory and transformer fundamentals (attention, GQA, RoPE / 3D Axial positional encoding, SwiGLU MLP) as applied to DiT.
- Comfortable working in a Linux environment with GPU compute, and reasonable familiarity with cloud or on-prem GPU clusters.
Preferred / Nice to Have
The following are not required for the role. Candidates with exposure will be prioritized.
- Prior experience in VFX, gaming, animation, or creative-tools industries.
- Hands-on experience with modern training backends: layer-specific LoRA/LoKr training, block-wise optimization, fp8 scaling, low-VRAM full fine-tuning, fused backward passes, and alpha-mask training — regardless of the specific frontend or CLI wrapper used.
- Experience creating quantized model artifacts: GGUF conversion (Q4_K_M, Q5_K_M, Q6_K, Q8_0), FP8/INT8 calibration, and deployment via diffusers-based or ComfyUI-GGUF pipelines.
- Exposure to emerging world / physical-AI foundation models (e.g. NVIDIA Cosmos-class, Genie-class world generation) for simulation, previs, or synthetic data generation — not required for this role.
- Experience with ComfyUI custom node development or Gradio/Streamlit-based internal tools.
- Docker, Kubernetes, and general DevOps familiarity — we run infrastructure on Azure AKS, so prior AKS or similar managed-Kubernetes experience is useful but not required.
- Familiarity with 3D/VFX tooling (Houdini, Unreal Engine, Unity, Nuke, Blender) and standard 3D asset formats (glTF/USD/FBX).
- Contributions to open-source GenAI/diffusion projects.
Education
Bachelor’s or Master’s degree in Computer Science, Machine Learning, Data Science, or a related field (or equivalent practical experience demonstrated through open-source contributions, competition results, or production deployments).