TNG Technology Consulting extended NVIDIA’s Nemotron 3.5 Lightning model with vision capabilities using a simplified approach designed for limited computing resources. Rather than training a complete vision model, the team adapted an existing vision encoder by training only a small MLP adapter layer of 34 to 40 million parameters to project visual embeddings into the language model’s space. Using entry-level GPUs, they achieved 58% accuracy on medical imaging tasks and competitive results on the MMMU benchmark.
