> For the complete documentation index, see [llms.txt](https://docs.concurrence.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://docs.concurrence.com/platform-overview/evaluating-amigo.md).

# Evaluating Concurrence

A technical evaluation path through architecture, data, control, delivery, measurement, and deployment, with concrete evidence to examine.

Evaluate Concurrence against a defined operational outcome and the requirements your organization must meet to deploy it. Agree on the workflow, baseline, required capabilities, and acceptance criteria before the demonstration. Identify any requirement that would prevent adoption, such as an unavailable integration, insufficient access controls, or an unmet operating commitment.

Follow representative cases from source data through execution to the external outcome. Include incomplete work, failures, and human intervention. Record whether each result came from a fixture, a provisioned test deployment, or production; those settings establish different evidence.

Use the [Technical Evaluation Packet](/platform-overview/evaluation-packet.md) for a worked evidence example and downloadable worksheets. Its [reference deployment](/platform-overview/reference-deployment.md) and [capability matrix](/platform-overview/capability-availability.md) make architecture and provisioning decisions explicit.

## 1. Establish the System Model

Start with [What Concurrence Does](/platform-overview/what-amigo-does.md), [Core Concepts](/platform-overview/core-concepts.md), and [How It Works](/platform-overview/how-it-works.md). Be able to locate four things: the retained source evidence, the current read view, the selected workflow configuration, and the outcome record.

**Evidence to examine:** a representative architecture and workflow map with the actual systems, interfaces, and responsible teams identified. Distinguish logical components from where they are hosted. Confirm the deployment's [data residency](/platform-overview/data-residency.md) and access paths separately.

For customer-owned agent or cloud requirements, use [Cloud Ownership and Integration Boundaries](/platform-overview/cloud-ownership.md) to review the execution model, permitted data exchanges, and action handoffs. The application host, agent runtime, and software supplier can have different owners.

## 2. Test the Data the Workflow Depends On

Use the [World Model](/data/world-model.md) and [Connectors and EHR](/data/connectors-and-ehr.md) to inspect how supported inputs become usable context. Trace one observation from its source into the selected entity view. Check what happens when the source is late, an identifier is missing, or two sources disagree.

**Evidence to examine:** source attribution, effective time where exposed, the relevant projection or mapping rule, and the context actually available to the workflow. A source record and a resolved caller binding are separate steps. Generated memory should also be distinguishable from current source facts.

Confirm that the records needed for the next action can be joined correctly and are current enough for that decision. Identify who owns mapping changes, data-quality exceptions, and source outages.

### Evaluate Memory on Tasks That Need Continuity

Use the [Memory](/agent/memory.md) workflow to test a remembered preference, a later correction, and a detail omitted from the loaded model. Compare the same task with no historical memory, authorized source retrieval alone, the compact model alone, and the model plus recall. Use a fixture where the supporting evidence is known so retrieval failure can be distinguished from reasoning failure.

Inspect which subject and workspace were searched, whether the answer preserves the source's meaning and time, and how the workflow handles no matches or unavailable recall. Measure the time from new evidence to usable later-session context separately from the latency of a recall request. Better continuity should be visible in the completed task, not inferred from the amount of stored history.

Include evidence attributed to the wrong person, repeated imports of one source, an older detail outside the recall window, and a correction that arrives during reasoning. Score evidence coverage, supported claims, correction handling, and task outcomes separately. A retrieved record must support the answer the agent gives.

## 3. Inspect Intent and Enforcement Separately

Read the [Agent Core](/agent/agents.md), [Context Graphs](/agent/context-graphs.md), and [Runtime Safety](/operations-and-safety/runtime-safety.md) together. Identify what is authored guidance, what is enforced in code, what is observed without blocking, and what requires a human decision.

**Evidence to examine:** selected agent and graph versions, state objectives, eligible tools, authorization, validation, and relevant failure cases. Trace a missing prerequisite through the actual operation. A prompt saying “verify before booking” and a booking tool rejecting an unauthorized request establish different controls.

## 4. Follow the Action to Its Destination

Follow a write through the configured direct-integration or connector path. Where approval is enabled, identify the exact workflow being gated. Then inspect what confirms delivery or an external mutation.

**Evidence to examine:** the recorded request, any required decision, the operation's result, and supported target acknowledgement or read-back. Include a timeout or ambiguous-response case and identify who reconciles it. [How It Works](/platform-overview/how-it-works.md#measure-the-whole-workflow) gives a running example; [Review Queue](/data/review-queue.md) explains the separately enabled connector review path.

## 5. Exercise the Real Channel and Human Fallback

Choose the [channel](/channels/conversations.md) the deployment will use. A text simulation does not exercise speech recognition, live audio interruption, carrier delivery, or operator phone joining. Messaging also requires the applicable identity, consent, and provisioning work.

Complete the [Channel Program Readiness](/channels/program-readiness.md) review for the intended purpose, audience, and provider route. Keep technical channel testing separate from the evidence permitting that program.

**Evidence to examine:** a representative interaction on the intended channel, delivery status where available, the supported [operator handoff](/operations-and-safety/operators.md), and the fallback when a person or destination is unavailable. Include an interrupted or incomplete interaction as well as a successful one.

Test overlapping work explicitly: change the requested appointment date while a lookup is running, let results complete in a different order, and exercise the supported cancellation or takeover path. Inspect both what the tool completed and what the person received. Measure first meaningful output separately from filler, and reconcile uncertain external actions before retrying. [Reasoning Engine](/agent/reasoning-engine.md#concurrency) explains the relevant execution boundaries.

## 6. Define What the Evaluation Can Establish

Read [Testing and Evaluation](/testing/testing.md) alongside [Intelligence and Analytics](/intelligence-and-analytics/intelligence.md). Separate conversation quality, workflow correctness, integration success, and the organization's final outcome. They may require different evidence sources.

**Evidence to examine:** the case set, baseline, scoring definition, required artifacts, missing results, relevant segments, and external outcome data. Ask whether the test used representative data and exercised the channel and integrations that matter. A favorable score supports only the claim its inputs and definition can measure.

Cost and latency comparisons need the same care: use a comparable workflow and account for unresolved work, retries, discarded reasoning, retrieval, and human handling. Separate replay with recorded results from a fresh model run. Agree percentile targets, concurrency and burst assumptions, retrieval budgets, and covered recovery failures before a pilot. See [Cost and Latency Optimization](/platform-overview/cost-and-latency.md).

## 7. Demonstrate a Reviewed Change

Use [Deployment Model](/platform-overview/deployment-model.md) to follow a candidate from authoring to validation and release selection. Inspect which components are pinned, which inputs can still change, and when a running interaction loads its configuration.

**Evidence to examine:** a configuration comparison, the affected regression cases, release approval, observations from new runs, and a tested recovery procedure. Restoring a prior agent version does not undo an external action that already occurred.

## 8. Agree on Ownership and Integration Boundaries

The [Operating Model](/platform-overview/operating-model.md) identifies decisions the organization and implementation team must make. Confirm who handles source failures, patient exceptions, consent changes, external reconciliation, access changes, and release decisions.

**Evidence to examine:** named owners and the interfaces they will use. If a team wants its own application or reporting, review the supported APIs and provisioned data-access path, the datasets available, and their freshness and permissions. If portability matters, identify the actual exportable artifacts and the work required to rebuild integrations and controls elsewhere. Open interfaces do not make an entire deployment automatically portable.

Review security and operating commitments alongside the technical results. Confirm the deployment's access controls, retention and deletion process, assurance evidence, support coverage, incident escalation, recovery expectations, and commercial assumptions. The [evaluation packet](/platform-overview/evaluation-packet.md#complete-the-enterprise-review) identifies the decision record for each area.

## Evaluate Proposed Extensions Separately

If the evaluation includes [object memory](/agent/memory.md#design-direction-object-memory) or the [Universal Reasoning Harness](/agent/reasoning-engine.md#design-direction-universal-reasoning-harness), give that work its own scope, implementation owner, and acceptance criteria. Both are documented as design direction beyond the current paths.

For object memory, test changed identity mappings, incompatible definition revisions, corrections, and revoked access. For the proposed harness, test failed required selections, late advisory guidance, uncertain action submissions, and responsiveness to control requests during long reasoning. Resumed work should use current authority and valid context. Equivalent typed and transcribed inputs, including corrections, should not create duplicate requests.

Record these as requirements for the extension until implementation and deployment evidence establish the behavior. Keep them separate from results obtained on supported production paths.

## Leave with a Deployment Decision

Capture the reviewed workflow, the evidence examined, unresolved dependencies, and the conditions for proceeding. Keep current behavior, deployment-specific configuration, private previews, and proposed work distinct in that record.

The next step should be concrete: resolve a source mapping, test a missing failure path, provision a channel, assign an exception owner, or approve a bounded release. The [Developer Guide](https://docs.concurrence.com/developer-guide) is the next layer when the implementation team needs exact contracts.


---

# Agent Instructions
This documentation is published with GitBook. GitBook is the documentation platform designed so that both humans and AI agents can read, navigate, and reason over technical content effectively. Learn more at gitbook.com.

## Querying This Documentation
If you need additional information that is not directly available in this page, you can query the documentation dynamically by asking a question.

Perform an HTTP GET request on the current page URL with the `ask` query parameter, and the optional `goal` query parameter:

```
GET https://docs.concurrence.com/platform-overview/evaluating-amigo.md?ask=<question>&goal=<endgoal>
```

`ask` is the immediate question: it should be specific, self-contained, and written in natural language.
`goal` is optional and describes the broader end goal you are ultimately trying to accomplish on behalf of the user. GitBook uses it to tailor the answer towards what is most useful for that goal.

The response will contain a direct answer to the question and relevant excerpts and sources from the documentation.

Use this mechanism when the answer is not explicitly present in the current page, you need clarification or additional context, or you want to retrieve related documentation sections.
