Introducing REFLEXThe new decision model / 2026

Intelligence,measured in outcomes.

Meet REFLEX. It reads the situation, weighs the available actions, and makes a clear choice. Text, photographs, and browser views become decisions your system can put to work.

A decision is only as good as what happens next.
Decision sequence / 001Interactive simulation
ObservationFind an eligible item under $80
procurement / request 24-018
Replacement desk lightBudget $80
◒Arc desk lamp / graphiteStock 8$68 · eligible
◐Studio lamp / brassStock 4$94 · over budget
◉Task light / ivoryStock 0$72 · unavailable
Three available actionsAwaiting action
The idea

From what you see to what you do.

Give REFLEX a goal, the situation, and the actions available. It selects the next move so your system can act—and you can see what happened.

02 / INTERACTIVE DEMOINTERACTIVE SIMULATION

See the decision.
Then the difference.

Explore three recorded-style synthetic scenarios. Choose an action yourself to see the resulting state and the checks that matter.

Scripted simulation · No live model inference
The goal

Observation

Available actions
Choose an action to simulate

Observed environment

Resulting state + verification

Select one of the available actions to reveal the next state.
Controlled synthetic environmentNo live inference
03 / FROM INPUT TO IMPACTA DECISION YOU CAN FOLLOW

A choice is just
the beginning.

REFLEX connects the moment of decision to the result that follows. The next action is explicit, and its effect can be checked.

01 / ACT
Make a decision from the available actions.

A model selects a candidate action from a text, photo, or browser observation.

02 / STATE
Let the environment change.

A purchase reduces inventory; a booking appears on the calendar; a repair updates records.

03 / CHECK
Verify the task and its boundaries.

Checks look for the requested result, service recovery, and preservation of unrelated state.

04 / CONFIRM
See whether the goal was met.

The result shows whether the action achieved the goal and respected its constraints.

−1

Inventory changes after a valid purchase.

+1

A compliant booking appears on the calendar.

✓

A repair changes only intended records.

OK

A health check confirms service recovery.

×

An incorrect action fails a constraint check.

04 / REFLEX RESULTSPRELIMINARY INTERNAL EVALUATION

REFLEX, measured.
In full view.

One model, two controlled task sets. Both results below measure REFLEX, with the sample size and incomplete tasks shown for each set.

REFLEX / core tasks500 tasks
79.4%
397 / 500tasks completed
397 completed103 incomplete
REFLEX / broader tasks450 tasks
45.1%
203 / 450tasks completed
203 completed247 incomplete
What was measured

Task success for REFLEX on two separate controlled task sets. These preliminary results should be read separately.

What remains open

These are not independent industry benchmarks. Comparable Jev, OpenAI, and Anthropic scores on the same REFLEX tasks are not available yet; broader confirmation is underway.

Source: REFLEX internal evaluation summaries · 2026-09-29 · Independent verification pending
05 / MODEL COMPARISONVENDOR DOCUMENTATION · 2026-09-29

One clear decision.
A distinct approach.

REFLEX is focused on turning a goal and an observation into one action from a supplied set. OpenAI and Anthropic publish broad models that generate responses and use tools; Jev offers another typed decision interface. This table compares documented roles and interfaces, not task accuracy.

REFLEX, measured hereTwo controlled task sets

Both results are for REFLEX, with sample sizes and incomplete tasks shown above.

Peers, documentedPublished interfaces

Jev, OpenAI, and Anthropic rows describe vendor-published roles and outputs.

Comparison scopeFacts, not a ranking

Model specifications and REFLEX task results answer different questions. Shared-task scores are still needed for a performance ranking.

ModelPositioningDecision interfacePublished facts
TYPESAFE / SPECIALISTJev

Typed answers to caller-defined questions about supplied state.

Returns a choice, score, or yes/no answer within the specified answer space.

$0.042 / 1MVendor-listed input tokens · typed decisions with confidenceVendor source ↗
OPENAI / EFFICIENTGPT-6 Luna

Cost-sensitive, high-volume text and image work.

Generates text or structured output; an app can provide functions and tools.

$0.10 / 1MInput tokens · 1.05M context · text and image inputVendor source ↗
OPENAI / BALANCEDGPT-6 Sol

Reasoning and everyday workflows with text and image input.

Generates text or structured output; tool calls can drive a workflow.

$2.00 / 1MInput tokens · 1.05M context · text and image inputVendor source ↗
OPENAI / FRONTIERGPT-6 Astra

Complex reasoning, coding, research, and computer use.

Generates responses and can use tools, including computer use in an agent harness.

$10.00 / 1MInput tokens · 1.05M context · text and image inputVendor source ↗
ANTHROPIC / EFFICIENTClaude Haiku 4.5

Fast general-purpose text and image work.

Generates responses and uses application-provided tools.

$1.00 / 1MInput tokens · 200K context · text and image inputVendor source ↗
ANTHROPIC / BALANCEDClaude Sonnet 5.5

Fast, broad model for coding, knowledge work, and image understanding.

Generates responses and uses tools inside an application or agent workflow.

$2.00 / 1MInput tokens · 1M context · text and image inputVendor source ↗
ANTHROPIC / FRONTIERClaude Opus 5.5

Long-running agentic coding and knowledge work.

Generates responses and uses tools inside an application or agent workflow.

$4.00 / 1MInput tokens · 1M context · text and image inputVendor source ↗
How to read this

REFLEX’s free beta is capped by decisions per month; peer prices are vendor-published input-token rates as of 2026-09-29, so these are different billing units. Output tokens, tools, discounts, and long-context rules can change peer costs. REFLEX task success appears above; a performance ranking would require every model to run the same tasks under the same rules.

REFLEX PRICING / EARLY ACCESS

Start free.
Decide with REFLEX.

Approved beta users get one shared monthly quota across the playground and developer API. There is no billing or payment method to set up during the beta.

$0100 decisions per monthPaid API pricing is coming later. We’ll publish rates before paid access begins.Join the beta waitlist
06 / WHY REFLEXREFLEX

A new kind of
decision intelligence.

REFLEX is made for the moment an application needs to choose. It brings the goal, the current situation, and the available actions into one clear decision.

From a photo to a browser view to a text workflow, the interface stays simple: give REFLEX context and a set of choices; get back the next action.

TextUnderstand the task
PhotoRead visual context
BrowserInterpret interface state
OneClear action to take next
Where decisions matter
ProcurementSchedulingData repairService recoveryPhoto packingExpense approvalAccess requestsTicket triage

These are examples of controlled evaluation tasks. Availability and performance in a live workflow depend on the application and its verification steps.

07 / FOR DEVELOPERSCONCEPTUAL INTERFACE

A decision has
clear boundaries.

The model’s role is to select among supplied candidate actions. The environment executes the chosen action; a separate verifier checks the resulting state.

InputGoal + observation + available actions
OutputSelected action
EnvironmentExecutes the action and changes state
VerifierChecks success and preserved constraints
Conceptual JSON · Not a working API{ }
{
  "goal": "Choose an eligible item under $80",
  "observation": {
    "type": "browser_screenshot",
    "state": "Three catalog options visible"
  },
  "available_actions": [
    "select_arc_lamp",
    "select_studio_lamp",
    "select_task_light"
  ],
  "selected_action": "select_arc_lamp"
}

Illustration of the decision boundary, not an endpoint, SDK, or live response.

08 / EARLY ACCESSTHE SCOPE OF REFLEX

Built for a clear
next move.

REFLEX is in early access. The current public results come from controlled tasks, and the developer playground is limited to approved beta users.

01
Browser capability is candidate-action selection in controlled interfaces, not unrestricted browsing or typing.
02
Photographic diversity is limited; the current evidence should not be generalized to arbitrary images.
03
Multi-job workflows do not demonstrate arbitrary long-horizon autonomy.
04
Performance may vary outside the controlled tasks shown here.
05
Confidence and calibration require separate evaluation.
Request beta access

Get in line for REFLEX.

Leave your email and we’ll contact you if a beta spot opens. Joining the waitlist does not grant playground access.

We’ll use your email only to contact you about beta access.

Meet the next move with REFLEX.