Intelligence,measured in outcomes.
Meet REFLEX. It reads the situation, weighs the available actions, and makes a clear choice. Text, photographs, and browser views become decisions your system can put to work.
From what you see to what you do.
Give REFLEX a goal, the situation, and the actions available. It selects the next move so your system can act—and you can see what happened.
See the decision.
Then the difference.
Explore three recorded-style synthetic scenarios. Choose an action yourself to see the resulting state and the checks that matter.
Scripted simulation · No live model inferenceLog in to try the live model
Observed environment
Resulting state + verification
A choice is just
the beginning.
REFLEX connects the moment of decision to the result that follows. The next action is explicit, and its effect can be checked.
A model selects a candidate action from a text, photo, or browser observation.
A purchase reduces inventory; a booking appears on the calendar; a repair updates records.
Checks look for the requested result, service recovery, and preservation of unrelated state.
The result shows whether the action achieved the goal and respected its constraints.
Inventory changes after a valid purchase.
A compliant booking appears on the calendar.
A repair changes only intended records.
A health check confirms service recovery.
An incorrect action fails a constraint check.
REFLEX, measured.
In full view.
One model, two controlled task sets. Both results below measure REFLEX, with the sample size and incomplete tasks shown for each set.
Task success for REFLEX on two separate controlled task sets. These preliminary results should be read separately.
These are not independent industry benchmarks. Comparable Jev, OpenAI, and Anthropic scores on the same REFLEX tasks are not available yet; broader confirmation is underway.
One clear decision.
A distinct approach.
REFLEX is focused on turning a goal and an observation into one action from a supplied set. OpenAI and Anthropic publish broad models that generate responses and use tools; Jev offers another typed decision interface. This table compares documented roles and interfaces, not task accuracy.
Both results are for REFLEX, with sample sizes and incomplete tasks shown above.
Jev, OpenAI, and Anthropic rows describe vendor-published roles and outputs.
Model specifications and REFLEX task results answer different questions. Shared-task scores are still needed for a performance ranking.
A focused decision model for goals, observations, and supplied actions.
Returns one action ID. The environment executes it; a separate verifier checks the result.
Typed answers to caller-defined questions about supplied state.
Returns a choice, score, or yes/no answer within the specified answer space.
Cost-sensitive, high-volume text and image work.
Generates text or structured output; an app can provide functions and tools.
Reasoning and everyday workflows with text and image input.
Generates text or structured output; tool calls can drive a workflow.
Complex reasoning, coding, research, and computer use.
Generates responses and can use tools, including computer use in an agent harness.
Fast general-purpose text and image work.
Generates responses and uses application-provided tools.
Fast, broad model for coding, knowledge work, and image understanding.
Generates responses and uses tools inside an application or agent workflow.
Long-running agentic coding and knowledge work.
Generates responses and uses tools inside an application or agent workflow.
REFLEX’s free beta is capped by decisions per month; peer prices are vendor-published input-token rates as of 2026-09-29, so these are different billing units. Output tokens, tools, discounts, and long-context rules can change peer costs. REFLEX task success appears above; a performance ranking would require every model to run the same tasks under the same rules.
Start free.
Decide with REFLEX.
Approved beta users get one shared monthly quota across the playground and developer API. There is no billing or payment method to set up during the beta.
A new kind of
decision intelligence.
REFLEX is made for the moment an application needs to choose. It brings the goal, the current situation, and the available actions into one clear decision.
From a photo to a browser view to a text workflow, the interface stays simple: give REFLEX context and a set of choices; get back the next action.
These are examples of controlled evaluation tasks. Availability and performance in a live workflow depend on the application and its verification steps.
A decision has
clear boundaries.
The model’s role is to select among supplied candidate actions. The environment executes the chosen action; a separate verifier checks the resulting state.
{
"goal": "Choose an eligible item under $80",
"observation": {
"type": "browser_screenshot",
"state": "Three catalog options visible"
},
"available_actions": [
"select_arc_lamp",
"select_studio_lamp",
"select_task_light"
],
"selected_action": "select_arc_lamp"
}Illustration of the decision boundary, not an endpoint, SDK, or live response.
Built for a clear
next move.
REFLEX is in early access. The current public results come from controlled tasks, and the developer playground is limited to approved beta users.
Get in line for REFLEX.
Leave your email and we’ll contact you if a beta spot opens. Joining the waitlist does not grant playground access.