It says “done.” The state disagrees.
The agent clicked Save. The right record, role, or workspace never changed. You need a checkable outcome.
Capture the browser path. Check the agreed outcome. Keep the failure as a case to rerun after your next model or application change.
A one-task pilot for teams building browser agents.
Start in an approved sandbox or test account.
Editor Viewer
“Role updated to Viewer.”
Still Editor after reload.
Start as Editor
Select Viewer → Save
Reload → still Editor First divergence
Agent engineering
Evaluation & reliability
Tool-use & post-training
When the final state is wrong, the useful question is what you can verify, reproduce, and test again.
The agent clicked Save. The right record, role, or workspace never changed. You need a checkable outcome.
A different session, permission, or starting state can hide the problem. The conditions matter as much as the clicks.
The trace gets reviewed. The incident gets closed. The same failure returns after the next model or application change.
Explore a role change that looks successful until the page reloads. Each stage adds evidence your team can act on.
After a reload, the member still has the Editor role. The checked page state disagrees with the agent’s report.
An illustration of the proposed pilot deliverable, using synthetic data.
Download sample JSONPass, fail, or incomplete. When the agreed outcome cannot be checked, the verdict stays incomplete. A success message alone is not proof.
Build on the runner and evidence you already have. Scope one outcome before you expand the evaluation set.
The pilot adds a scoped capture and checking workflow around it.
The artifact connects that path to starting conditions and a checked result.
The failure becomes a case to rerun against a new model, prompt, policy, or app.
The pilot preserves the first meaningful divergence for investigation. Automatic root-cause diagnosis and repair are outside its scope.
A focused pilot for teams that struggle to reproduce or grade a recurring browser-agent failure.
Agree on the starting state, allowed actions, success condition, and model and environment versions.
Use a sandbox, synthetic workflow, or test account. Agree on access, redaction, retention, and permitted use first.
Ordered actions, checked outcome, divergence evidence, and reproduction instructions with known reset requirements.
Pilot scope and acceptance criteria are agreed before work begins. Production customer data is not required.
We’ll email you about pilot availability and next steps. If there’s a fit, we’ll scope one failing browser task, its starting conditions, and an observable success condition with your team before any run. Joining the waitlist does not reserve or guarantee a pilot place.
The pilot is designed to work around your existing agent runner. Samelogic’s role is capture, evidence, outcome checking, and regression packaging. The runner and environment are agreed when we scope the task.
No. The pilot identifies the first meaningful divergence and preserves evidence for your engineering or evaluation team to investigate. It does not promise automatic diagnosis or repair.
The verdict is incomplete. A click, a success message, or the agent’s own report is not enough to prove a task finished. We agree on observable conditions and the evidence your approved environment can provide before running the task.
No. Start with a synthetic workflow, test account, or approved sandbox. We scope access, retention, redaction, deployment requirements, and permitted use before any run.
If the artifact proves useful for investigation or catches a regression, we can scope the next task family. Acceptance criteria and commercial scope are agreed before custom work begins.
Join the waitlist for the one-task pilot. We’re bringing teams on gradually, starting with a concrete browser task and an agreed way to check it.