VLM Run Gateway: Run GLM-OCR, DeepSeek-OCR-2, and dots.mocr with an OpenAI-Compatible API
VLM Run Gateway is a unified platform offering a single OpenAI-compatible endpoint for open-weight OCR and vision-language models, including DeepSeek-OCR-2,…
VLM Run Gateway is a unified platform offering a single OpenAI-compatible endpoint for open-weight OCR and vision-language models, including DeepSeek-OCR-2,…
mLateOn is a multilingual retrieval model that combines a multilingual mmBERT encoder with ColBERT-style late-interaction token matching. Evaluated on the…
The Open Discovery Challenge introduces a computational leaderboard for scoring AI-designed malaria drug candidates across six dimensions: whole-cell activity, target…
Liquid AI introduced LFM2.5-VL-3B, a vision-language model that pairs a SigLIP2 400M NaFlex vision encoder with the LFM2.5-2.6B text backbone.…
FINAL-Bench released LEADBOARD, a drug-property-prediction benchmark spanning 21 boards and more than 18,000 held-out test compounds. The team found that…
Google DeepMind announced a research partnership with Fenris Creations, the studio behind EVE Online, to use the game’s persistent virtual…
Researchers at Google introduce EnvHarness to address the problem that agent learning environments are static and don’t adapt to an…
Researchers introduced SWE-bench Science, a repository-level benchmark comprising 119 tasks drawn from 98 GitHub repositories across 20 scientific domains, to…
Researchers introduce MemTrapBench, a benchmark showing that retrieved memories can actually harm large language model task performance even when the…
Tencent researchers present SkillEvo, addressing the problem that AI agent skills typically fail to improve from interaction failures beyond a…
Researchers at Shanghai Jiao Tong University present Repo0, a system for generating complete software projects with proper modular architecture directly…
Z.ai’s GLM-5.3 model was trained with post-training scaling improvements aimed at vulnerability detection, achieving 84.5% on the CyberGym benchmark for…