SpanCalib-VLM combines a multimodal sequence tagger with a fine-tuned generative vision-language model to detect hallucinations in large VLMs, using a Union-Calibrated Fusion strategy that rescores candidate spans from the generative model with calibrated probabilities from the sequence tagger. On the SHROOM-Visions benchmark the ensemble reaches 0.41 Pearson calibration correlation and 70.7% detection accuracy.