Models

AutoTrust JEV-27B-VL Beats GPT-4o in Image Scoring

AutoTrust has launched JEV-27B-VL, an open-source vision-language model that evaluates images and text in a single forward pass to deliver highly calibrated decisions for autonomous agents.

AlphaSignal2 days agoModels
Image: AlphaSignal

AutoTrust has introduced JEV-27B-VL, a new vision-language model designed to make rapid, structured decisions from visual and textual inputs. Built on the Qwen3.8-27B backbone under an Apache-2.0 license, the model integrates a LoRA adapter and a specialized decision head. This architecture allows the model to bypass slow, token-by-token label generation. Instead, its System 1 interface returns calibrated probabilities for yes-or-no questions, ratings from 0 to 5, or selections among 2 to 256 choices in a single forward pass.

The model has demonstrated strong performance across several benchmarks, surpassing prominent proprietary models. JEV-27B-VL achieved a 78.3% score on VL-RewardBench and 73.2% on Plan-RewardBench, outperforming GPT-5 and Gemini-3-Flash in agent judging tasks. In practical control loops, the model achieved a 75% success rate in MuJoCo robot pick-and-place simulations and a 95% success rate across 60 multi-step browser automation tasks. Additionally, in a zero-shot video recommendation test using only cover images, it nearly matched a collaborative filtering system trained on 59,045 users, scoring an AUC of 0.727 compared to the traditional system's 0.728.

For machine learning practitioners, JEV-27B-VL offers a highly efficient alternative for classification, ranking, routing, and evaluation tasks. The model supports context windows up to 256K tokens and is deployed via a POST /v1/decide endpoint built on top of vLLM. Because the decision head outputs calibrated probabilities, developers can establish reliable confidence thresholds. This allows them to route routine, low-latency decisions through the fast System 1 process, while reserving more complex, open-ended queries for System 2 reasoning. The model's underlying text decision capabilities have already proved popular, with its Hugging Face page recording over one million downloads.

This is our own summary of reporting by AlphaSignal

More in Models