Give the server a task with an answer you already know. That first check establishes whether the assistant reaches the intended system and receives useful information.

Then exercise a missing record, a denied action, and a relevant dependency failure. Those cases help you decide whether the service can support your workflow and which limitations you would have to operate around.

An evaluation sheet for a real workflow

For each required action, record the following:

Question Evidence
Does the tool exist? Tool name, description and required inputs
Does it reach the right data? Known input compared with the source system
Is access enforced? A denied request from an unauthorized identity
Is the result usable? Clear success, empty-result and error responses
Can a write be retried? Documented behavior and a controlled duplicate test
Can you investigate a failure? The record or log available to the operator
Can you stop access? Revocation procedure and observed result

A description of a safeguard is not a test result. Keep documented behavior and observed behavior in separate columns.

Test through the client you intend to use

MCP standardizes communication, while supported features vary across implementations and versions. The architecture documentation describes the relevant protocol concepts. Your chosen combination still needs an end-to-end check.

Connect the actual client, authenticate as the intended user, and complete the task. A request sent directly to the server can help diagnose problems, but it may skip client approval, tool selection, or authentication behavior that matters in daily use.

Keep the first evaluation reversible. For a write, use synthetic records and a clearly named test destination. If a timeout occurs, inspect the destination before issuing another action.

Make a decision from the failures

Consider a hypothetical support-case server. It finds a known case correctly, returns a clear error for a missing case, but can read a case belonging to an unauthorized team. Record all three outcomes. The successful lookup does not offset the access failure; the server is unsuitable for that shared-data workflow until the boundary is fixed and retested.

A different candidate might enforce access but omit an optional attachment preview. That can be an acceptable limitation if the user can open the source system and the extra step is clear. The decision depends on the job the team needs done.

Agree on mandatory requirements before the trial. Separate stop conditions, workable limitations, and improvements that can wait. Ask the operator to explain how a failed action is investigated using evidence the customer can actually obtain.

The evaluation record should end in a decision: suitable for the named workflow, suitable with stated limits, or not ready for that use. Include the unresolved behavior and the evidence needed to revisit it. A list of green feature names without that decision leaves the buyer to interpret the risk later.

Measure completed work

If time savings are part of the purchase decision, measure a representative task and record review and correction work too. Avoid comparing an automated successful run with a manual workflow that includes all the exceptions.

A useful result can be modest: the assistant found the correct record, prepared a valid draft, and left the final action to a person. Record that outcome accurately instead of calling the whole process autonomous.

We expose OBTO application operations through MCP. Our connection guide gets you to the first read. For an evaluation, compare the returned application and artifact with the destination you intended to inspect.