EvoGenUI-Bench: Evaluating LLMs as Multi-Turn Generative UI Assistants
EvoGenUI-Bench evaluates LLMs’ ability to maintain and evolve interactive web interfaces across multiple turns, comprising 150 five-turn tasks across three…
EvoGenUI-Bench evaluates LLMs’ ability to maintain and evolve interactive web interfaces across multiple turns, comprising 150 five-turn tasks across three…
SafeAtlas-VL is a dataset of 1.5 million annotated instances for multimodal safety evaluation that replaces binary classification with a five-level…
BLARM is a feed-forward method for generating animated 3D meshes from monocular video without explicit rigging or skeletal annotations, representing…
This paper studies whether function vectors — task-specific latent directions derived from in-context demonstrations — transfer across languages for emotion…
Chat-Edit-3D++ is a dialogue-based approach for editing 3D and 4D scenes that uses an LLM to interpret flexible text input…
RECAP-Forcing is a training-free inference method for long video generation that organizes memory by content novelty rather than temporal recency,…
This paper introduces Data Intrinsic Consistency (DIC), a self-scoring metric measuring sample-level inter-component consistency through Visual Information Consistency and Response…
SpanCalib-VLM combines a multimodal sequence tagger with a fine-tuned generative vision-language model to detect hallucinations in large VLMs, using a…
Dynamic Important Example Mining (DIEM) adaptively selects and weights training data throughout reinforcement fine-tuning, combining a gradient-alignment importance estimator that…
This paper augments the Aardvark Weather deterministic end-to-end AI forecasting model with probabilistic capability by adding learned, input-dependent noise at…
This paper presents a taxonomy for analyzing multilingual multimodal misinformation on social media, built from a large-scale seven-language Twitter/X dataset…
VLAct is a vision-language-action model using representation-centric continued pre-training on diverse robot data, applying VLM-prior preservation, multi-head action co-supervision, and…