# AI Employees: genuine isolated model execution

24 September 2026. A real worker completed one fictional inline checklist through the application, using its own employee sign-in and a real Claude Sonnet model session. This is separate from the [manual control rehearsal](ai-employees-manual-lifecycle.md).

## Observed result

Job `6ab5a2817393e89fda21a4df`, **Actual model: review an inline fictional checklist**, reached **done**. The runtime reported model cost **$0.0350992**. Its persisted answer:

> The missing check is the contact form submit test (item 2), which matters because an untested form may silently fail to deliver visitors' messages, so launching would risk losing leads without anyone noticing.

The actual inline input was:

> This is an isolated software acceptance test. All necessary information is here: Fictional launch checklist: (1) page title is set, (2) contact form submit has not been tested, (3) mobile layout was checked. Identify the one missing check and explain why it matters in one sentence. Do not read any files, APIs, docs, chats or records. Do not change anything or contact anyone. Use ./ae step note to record your reasoning, then ./ae complete with your final answer. No approvals or external action are needed.

The runtime signed in with this employee's password and authenticator, sent hello, fetched its briefing, leased the job, invoked the model, recorded its reasoning and completed the work. The model initially omitted the `note` argument from its step command; that call did not create a valid step. It corrected the invocation and completed successfully. This was not replaced by a hand-written result.

[Sanitized runtime and persisted job evidence](assets/ai-employees/actual-model-proof.json).

## Isolation and cleanup

- Source was copied to `/tmp/ai-employee-actual-review`, with an initially empty data directory. The original `projects/ai-employee/data` was neither loaded nor modified.
- Local provider port3414 initially reported zero agents. Exactly the dedicated training employee was provisioned with `start:false`, verified stopped, and explicitly started.
- A loopback proxy on3314 allowed only this employee's authentication, briefing, hello, lease and exact job step/complete/fail routes. Business APIs and chat were refused.
- The rehearsal limited available Claude tools to Bash and the `ae` commands to step/complete/fail, with a $0.50 session cap. The Bash-only availability setting was subsequently adopted in canonical source and regression-tested; the exact command allowlist and cost cap remain review-specific. No additional host tools were enabled.
- No production/customer records, attachments, messaging, external business action or payment was involved. The model service itself was genuinely invoked.
- After completion, the provider was stopped, the employee paused and its application budget returned to zero. Existing workers and main API3312 were untouched.

The original source's `--allowedTools` setting was an automatic-permission list, not by itself an exclusive list of available tools. Canonical source now explicitly supplies `--tools Bash`, matching the genuine rehearsals and its documented worker contract. This still does not constitute a complete filesystem or network sandbox; see the [source repair and regression](ai-employee-step-command.md).

## Limits

This proves a real, bounded model task and persisted result. It does not prove customer interaction, external integrations, periodic scheduling or every template's ability to perform business work. File processing remains separate from the exact-record rehearsal below.

## Real model approval, rejection and resume

The owner created job `6ab5a3447393e89fda21a4e7` through **Give work → Send** in Studio, with no attachments. Its instruction required the model to ask whether it could record a fictional checklist as approved, stop on WAIT, and respect the owner's decision without any external action.

The real model invoked `./ae approval other 0 "May I record the fictional launch checklist as approved?"`. The application entered **waiting_approval**, and the model ended its turn. The owner opened the actual approval drawer and selected **Reject**. That drawer had no decision-note field; no invented note was supplied through the UI.

On its next poll, the runtime leased the job again and resumed the saved model session with the rejection. The model completed with:

> The owner rejected the fictional proposal to record the launch checklist as approved. I did not approve anything and performed no action.

The persisted total model cost was **$0.0479462**. While waiting, the runtime retained the first turn's cost locally; it posted the accumulated cost on completion. The provider was then stopped and the employee paused with zero application budget.

[Actual model approval/resume proof](assets/ai-employees/actual-model-approval-proof.json). The isolated command/endpoint allowlists were extended only for this employee's approval request; no business action was enabled.

## Exact fictional record attachment — passed

The schema catalog has no standalone note datatype. A fictional **Task** was therefore created with a note field, no assignee and no workflow/customer relationship: `6ab5a4287393e89fda21a4f3`, **Fictional launch checklist attachment**. Its note contains a verification marker that was omitted from the job prompt and attachment summary.

The owner used **Give work → Records → Task**, searched the title, selected the exact row and attached it. Job `6ab5a4537393e89fda21a4f4` was sent through the actual UI. The model can retrieve only `GET /repository/get/task/6ab5a4287393e89fda21a4f3` through the review proxy.

Before execution, the employee's own token received **200** from the canonical exact-record read, and **403** from a canonical same-value `update-partial` request against that same fictional task. This is an actual AppEngine write refusal for the tested identity and route, distinct from the review proxy's restrictions; it is not a claim about all possible routes or groups.

The genuine model fetched the exact record and completed with **LANTERN-4827**, identified the missing contact-form submit test and reported that it changed nothing. The marker appeared only in the task's note, not the job prompt or attachment summary. Persisted cost: **$0.028183**. Provider stopped and employee paused with budget zero afterward. [Record proof](assets/ai-employees/actual-model-record-proof.json).

The fixture was created through the owner API, not an invented UI creation flow:

```http
POST /repository/create
Content-Type: application/json
orgid: <your training organization>
Authorization: Bearer <owner token>
```

```json
{
  "datatype": "task",
  "isNew": true,
  "data": {
    "name": "ai-attachment-training-20260924-valid",
    "title": "Fictional launch checklist attachment",
    "status": "new",
    "description": "Harmless internal training record for exact-read AI attachment acceptance. Not a customer record.",
    "note": "Verification marker: LANTERN-4827. Fictional checklist: page title checked; mobile layout checked; contact form submit test is missing. No customer or external operation is involved.",
    "assignTo": []
  }
}
```

`isNew: true` is required by this repository create API. An initial request without it was refused; it was not counted as successful creation.

## Actual Studio-managed Provision and Check provider

A second isolated manager on3415 began with an empty data directory and zero agents. Only the review API process was configured with its URL/key; original managers and global AI defaults were unchanged. The employee was paused, budget zero, and had no open work.

The Edit drawer did not expose a runtime-provider selector. The tested prerequisite was an owner API update:

```http
PUT /ai-employees/tutorial-lifecycle-20260924
```

```json
{"runtime":{"provider":"session-manager","model":"sonnet"}}
```

In the actual Studio **Connection** view, the owner clicked **Provision**, then **Check provider**. Provision created the agent on the isolated manager and persisted its external ID. Check provider reported that the agent existed and its process was running, with detail **locked: Its login is locked — switch it on in AI Employees.** The employee itself remained **Paused**; no job or model session ran.

This is an important distinction: a provider process can be running while the employee is locked and cannot take work. The isolated provider agent was explicitly stopped after capture; its current job was null. [Managed provider proof](assets/ai-employees/managed-provider-proof.json).

![Managed provider before provisioning](assets/ai-employees-local/managed-before-provision.png)

![Real provider check with paused employee](assets/ai-employees-local/managed-provider-checked.png)

## Genuine private-file processing and fixed step command

The owner selected **Private** in the real Upload control, uploaded `fictional-launch-checklist.txt`, attached it and sent job `6ab5a755b71e20a3a9111a62`. The UI defaults to **Public**; selecting Private was an explicit step. The training document contained only:

```text
FICTIONAL TRAINING DOCUMENT — no customer information
Verification marker: MAPLE-7392
Launch checks: page title checked; mobile layout checked; contact form submit test is missing.
This document is only an application tutorial acceptance fixture. Do not contact anyone or modify any business record.
```

The marker was omitted from the work instructions. After the owner switched the employee on in Studio, the isolated managed runtime started and fetched only this exact file through `POST /repository/file`, using its own employee identity and `{location: "learner-recovery-mubo2g1c/presentation/fictional-launch-checklist.txt", encoding: "utf8"}`. The review proxy and CLI refused other file/resource routes.

The genuine model returned **MAPLE-7392**, correctly identified the missing contact-form submit test, and persisted **Reviewed the private training document** as a note. This also verifies the [step command source fix](ai-employee-step-command.md) on a real model turn. The job reached **done** with reported cost **$0.035042**.

The managed provider was stopped, the employee paused and its budget restored to zero. Both its existing token and a fresh sign-in returned **403** afterward. [Model/file proof](assets/ai-employees/actual-model-file-proof.json), [final private-file/authentication checks](assets/ai-employees/final-managed-private-checks.json).

The upload's stored URL was initially unsigned and returned403. That exposed a separate attachment-open UI defect: private documents required a fresh authorized link. The composer and work drawer were updated to request a five-minute signed URL for the stored path whenever opened, without saving that signature in the work item. A fresh signed read returned200; the unsigned read remained403. Actual browser click verification is recorded in the separate attachment repair report.
