Dots: ongoing work in ChatGPT
OpenAI presents an always-on agent with its own cloud computer, connected apps and remembered context. Give it an ongoing goal and review the work and decisions it brings back.
Keep the work alive across sessions. Give the organization control over who can act, which tenant they serve, when humans intervene and whether the workflow pays back.
Four vendor promos make the landscape tangible: agents with memory, tools and work that continues between conversations. These are marketing demonstrations, not independent evidence of reliability, security or ROI.
OpenAI presents an always-on agent with its own cloud computer, connected apps and remembered context. Give it an ongoing goal and review the work and decisions it brings back.
Meta presents a persistent cloud computer and browser for tasks across connected services. The pitch includes background progress, remembered goals and approval requests for sensitive actions.
Grok presents bots with their own computers that work across apps and inboxes, remember conversations and return for approvals. Its examples span sales, operations and engineering.
Microsoft's broader Copilot promo includes Autopilot, a persistent agent with its own identity, memory and workspace in the organization's tenant. Autopilot is expanding through a private preview.
A productive persistent agent needs tools, working systems, tests and a way to learn from failure. That habitat can support software maintenance, an inbox or an operational workflow.
An agent's natural habitat is an environment it can actually use: a repository, a mailbox, internal records, a browser or application, test data and a clear definition of done. Persistence gives it time to complete the loop. The organization supplies the scope, safeguards and visible delegation that make the result worth accepting.
Imagine a bot assigned to maintain an internal application. A permitted error feed reports that a form sometimes loses a saved value. The bot claims the issue, starts a clean workspace at a recorded commit, loads synthetic test accounts and reproduces the failure in the application. It keeps the failing steps and logs alongside the task.
Then it reads the relevant code, writes a regression test that fails for the original defect, makes the change and runs the relevant checks. It opens the application again, follows the same journey and checks nearby behavior: keyboard navigation, a narrow screen, an expired session or a failed request, as applicable. A screenshot or recording makes the visible result inspectable; assertions establish what the test actually proved.
The output is a reviewable change with the reproduction, diff, test results, remaining uncertainty and rollback path. If a release needs approval, the bot prepares that decision and can move to other authorized work while it waits. After an approved release, it checks the agreed health signals and reports whether the original symptom returned. That is a complete maintenance loop.
Vendor example: Cursor's cloud-environment account describes simplifying build commands and giving agents computer use and recordings to validate running software. It supports the environment pattern; the working day above is our proposed design.
The same operating pattern can take over work previously assigned to a person: tracking deliveries, maintaining support records, reconciling routine documents or preparing recurring reports. A coordinator's day often connects messages, decisions and updates across several systems. An agent can perform that chain when the inputs, permitted actions, exceptions and completion evidence are explicit. How much of a human role it can replace is a question for measured outcomes, including the supervision and recovery it still needs.
OpenSpec gives us a useful starting method. Its documented purpose is managing software specifications, with proposal, requirements, design and task artifacts that guide implementation and verification. Our proposed extension is to apply that discipline to business workflows: describe the trigger, trusted input sources, records it may read or change, people it may contact, deadlines, acceptance criteria and escalation path. The specification loop already offers a way to make these decisions reviewable.
For vendor coordination, the specification might permit updating a delivery estimate for an existing purchase order and sending a routine acknowledgment. It can reserve changed prices, payment details, contract terms and disputed deliveries for a human. Version that contract with the workflow and its tests. OpenSpec can organize the requirements; the deployed application's authorization checks must enforce them.
Hermes documents an email gateway that receives messages and replies through IMAP and SMTP. It also distinguishes that gateway from its Himalaya email skill for inspecting and managing mailbox messages. Its MCP support provides a route to additional tools. These components are a starting point for a workflow with internal systems; the integration and business rules still have to be built and tested.
Imagine a dedicated vendor mailbox. A supplier reports that an existing order will arrive two days later. The workflow checks the sender and order against trusted records, extracts the proposed date and prepares an update. A narrowly scoped service checks the allowed fields and current record version, applies the authorized change and returns a receipt. Only then does the agent send the permitted acknowledgment to the approved recipients. If the update fails or the order cannot be matched, it escalates instead of claiming success.
Keep a durable link between the incoming message, internal update and outgoing reply. Use a stable operation key to prevent a duplicate message or restart from repeating the write; track reply delivery separately so a failed send can recover without reapplying the update. Reconcile uncertain delivery status before resending. External correspondence remains task data. A vendor's message must not become permission to run commands, change the workflow or access another customer's records. The gateway's sender controls are useful, but a vendor who may supply an update should not inherit an operator's tool authority.
Begin with synthetic mail, a test mailbox, disposable records and an isolated worker. Limit filesystem access, host mounts, credentials and outbound destinations to the task. Keep each tenant's memory, files and connections separate. Start with drafts and proposed changes; enable a bounded write or reply only after its acceptance and failure cases pass.
Mailbox reads, internal writes and external sends need separate grants. Keep credentials in a broker or narrowly scoped connector, and validate the target, permitted fields, recipients, attachments and spend at execution. A container can constrain local commands while a connected mail service still sends a harmful message with valid credentials. The tool boundary has to govern that action too. The Hermes security guide documents container backends, command approvals and other configurable defenses; their presence still requires an appropriate deployment and review.
Test hostile instructions in emails and attachments, forged senders, duplicate messages, stale records, changed recipients, partial writes and unavailable services. Require the agent to preserve its original authority throughout. Store the evidence needed to reconstruct the action, restrict access and avoid retaining entire private conversations by default. Use the decision ledger and personal-data guide when defining the deployment's records and retention.
The organization should know which workflow an agent operates, whose authority it uses and who can stop it. People receiving its correspondence should know they are dealing with an automated assistant and how to reach the accountable human. Give the agent an honest identity and signature, and explicit limits on promises, meetings and commitments. Put the deployment through the organization's normal ownership and budget process.
There is a funny, awkward illustration in this Reddit account of a Hermes vendor coordinator. The author says they quietly substituted an agent for a planned hire, gave it a fictional human identity on company email and let it coordinate with suppliers. A vendor then invited the supposed colleague for coffee, and the agent agreed to find a time. This is an unverified personal account, rather than evidence of productivity or typical Hermes behavior.
The invitation makes the missing boundary easy to see. A correspondence workflow can create expectations and commitments even when it changes no code or money. An honest assistant identity, a named owner and a rule that routes personal invitations or new commitments to a human would make that moment manageable. Useful automation earns trust through visible, authorized work and recoverable mistakes.
For a bot doing sustained coding, research, coordination, tool use and verification, substantial token throughput should be an ordinary operating consideration. Every iteration brings context back into the model: code, documents, correspondence, instructions, tool results and earlier decisions. Running several bounded tasks over a week can produce a very large usage total.
One billion tokens per week is an editorial planning scenario here. For example, 10,000 model calls averaging 100,000 combined input and output tokens total one billion tokens. That arithmetic includes repeatedly supplied context and any cached input counted by the provider. It does not describe a billion tokens of newly written code or establish a measured industry norm.
Budget fresh input, cache reads, cache writes, reasoning and output according to the provider's actual accounting, alongside tool and machine costs. Then divide total operating cost by accepted tasks. A bot that fixes valuable defects or completes useful business transactions may justify substantial usage; a bot that keeps rereading its context without producing accepted work needs a different loop. Give each task a spend ceiling, bounded retries and a stop when repeated attempts produce no new evidence.
An agent workspace does not have to reproduce a human's desk or working day. Make the environment easy to discover and cheap to reset. Provide documented routes to start the application or workflow, run its checks and restore realistic fixtures, with clear output when a service fails. Supervise long-running processes outside the model so the bot can concentrate on the task.
Prefer a stable command or API for precise operations, and provide browser or native application access wherever behavior must be experienced. A passing unit test cannot show that a dialog covers the save button. A screenshot cannot prove that a permission boundary holds. The agent needs the evidence appropriate to each claim, including failure cases the implementation did not anticipate.
Keep the task queue, acceptance criteria, checkpoints and decisions in durable storage. Record which code, workflow version and environment produced each result. Resume by inspecting current state, including completed writes and expired approvals, so a restart cannot quietly repeat a consequential action.
Our first implementation would have a tenant-scoped task store and scheduler, an isolated worker running the chosen harness, a tool broker enforcing named targets and short-lived credentials, an approval service and an evidence store. Give tasks explicit states such as queued, running, waiting for approval, verified, failed and completed. Atomically claim work with an expiring lease, and check permissions at the point of execution.
Begin with one worker and one recurring job, such as preparing fixes for a bounded class of bugs or handling routine delivery updates for one team. Add concurrency when the queue warrants it. A separate reviewer can challenge the acceptance evidence, while execution gates remain application policies owned by accountable humans. The one-hour decision policy and decision ledger below describe those boundaries.
Grok Bot through an eligible SuperGrok or Cursor plan is a useful place to experience this workflow. SpaceXAI's launch account describes internal engineering bots reproducing UI bugs and handing fixes to other bots. Its engineering marketplace offers templates to import, including a SWE bot by Cursor described as coding, issue triage and production monitoring. These are vendor-described capabilities.
Import a suitable template, inspect its instructions and adapt its workflow around a disposable repository and synthetic accounts. Test whether it can reproduce a bug, repair it, exercise the application and leave evidence another person can assess. Cursor's documented harness is a useful reference for this full loop. Keep the experiment portable by saving the task contract, checks and artifacts in formats another worker can use.
Availability checked 1 October 2026: the Grok Bot overview and FAQ lists eligible plans and weekly included usage with additional token billing. It says bots belonging to a user share a persistent computer, files, browser and logins. Separate bot names do not establish tenant isolation. An importable template does not establish free execution or a FOSS licence.
For a system the organization will operate, our preference is free and open-source software (FOSS) with replaceable model providers. The MIT-licensed OpenCode and Pi projects are starting points for the worker harness. Pair one with the durable coordinator and controls above, then evaluate the same tasks and acceptance criteria used in the commercial experiment.
For the operational example, the Hermes source repository provides an inspectable foundation for the persistent assistant, messaging and tool loop. Use OpenSpec to keep the workflow's requirements and acceptance criteria portable. Select the worker for the actual job, then test its connectors, permissions and recovery against the same contract.
A harness licence covers its software; model licences, inference bills, browser tooling and integrations need their own assessment. Isolation and authorization still belong in the deployed system. Keep the Pi permission boundary explicit, and compare model options through the local model and cloud model guides.
The first milestone is a bot that repeatedly closes one useful loop and leaves enough evidence to trust the result. Its natural habitat makes independent work possible. Accepted outcomes, recovery and the 90-180 day full-cost review determine how much production scope it earns.
A persistent agent retains task state and resumes after events, pauses or failures. A multi-tenant control plane owns identity, permissions, scheduling, approvals and evidence separately for each customer or organizational unit.
Bind authenticated users, workers, memory, retrieval, files and credentials to a tenant. Enforce this at every storage and tool boundary. A project folder or tenant ID written in a prompt does not provide isolation.
Persist tasks, checkpoints and approval state outside the model context. Use bounded retries, per-tenant queues and cost limits, expiring worker leases and idempotency keys for external writes. Resume against current policy after a restart.
Grant named tools, targets and budgets explicitly. Broker short-lived credentials outside prompts. Separate the component that approves an action from the worker executing it, and provide a tenant-level pause and credential-revocation path.
This is a proposed operating policy, not a verified feature shared by the tools below. Choose the eligible voters, quorum, vetoes and timeout behavior before starting work.
| Action | Team vote | No decision at deadline |
|---|---|---|
| Research, tests or draft in an isolated workspace | Optional inside the existing task grant | Continue only within its original scope and budget |
| Bounded staging change or internal ticket update | Allowed if the accountable owner delegated it in advance | Pause the write; prepare a draft or escalate |
| Production release, payment, deletion or wider access | Does not replace a required release, finance or security approver | Remain blocked pending the required approval |
The model may explain the options and recommend a next step. Deterministic application policy counts votes and grants authority. External documents and chat content remain untrusted inputs; they cannot rewrite that policy.
Customer feedback votes: OpenHeard lets users post and vote on requests, with roadmap, changelog, API and MCP routes. Those votes help prioritize demand. They do not authorize a production action or establish an internal Slack/Teams approval ballot. Its official site describes this feedback loop.
Log enough to reconstruct authority and outcome without copying every secret or private conversation into a permanent trace.
Tenant, task and request IDs; actor identity; model, tool and policy versions; payload hash; deadline; votes and vetoes; permission decision; execution result; spend and rollback reference.
Use append-only audit events with restricted access and tamper detection. Redact secrets, minimize personal data and set retention and deletion rules. Review the GDPR and data-processing guide for the actual deployment.
Test cross-tenant reads and writes, forged and replayed callbacks, duplicate votes, stale approvals, changed payloads, timer/reply races, retries after partial writes, provider outages and the emergency stop.
Compare the worker, task coordinator, execution policy and approval channel. Each layer has a different job and a different deployment boundary.
Sorted by role in the enterprise plan. These are complementary building blocks; the shortlist is an architectural assessment, not a tested integration.
Run persistent assistants inside a bounded task grant.
A persistent assistant and messaging gateway. OpenClaw's security guide assumes one trusted boundary per gateway; separate gateways and credentials for mixed-trust tenants.
Enterprise fit: Fit as execution workers behind your own policy, credentials and task ledger. OpenClaw requires a separate trust boundary for mixed-trust tenants; neither popularity nor a command-approval feature establishes the delegated ballot.
Hermes: 100,000+ stars; repository created July 2025; published GitHub release. OpenClaw: 100,000+ stars; repository created November 2025; published GitHub release.
Put bounded coding tasks behind a durable, reviewed queue.
Bounded code tasks behind a reviewed queue. Your coordinator owns tenancy, checkpoints, identity, approvals and retries. Select it on required controls and total cost.
Enterprise fit: Strong fit when the coordinator implements tenant isolation, a durable decision ledger and tool authorization. The maturity of a coding harness does not transfer automatically to your custom coordinator.
No single repository or popularity score: assess the selected harness and coordinator separately.
Inspect the gateway that decides and records tool actions.
Give coworkers separate browsers and files behind a gateway that documents policy evaluation and audit recording before execution. This alpha defaults to a single administrator without sign-in. Configure authentication before sharing it and keep lower-level computer endpoints private.
Enterprise fit: The most directly aligned action-policy and audit reference here. Pilot authenticated gateway enforcement and isolated computers, then test tenant boundaries, durable decisions and bypass paths. The default single-admin mode must be disabled.
1,000-10,000 stars; repository created August 2026; published GitHub release.
Adapt resumable approval cards and write interrupts in team chat.
The Channels SDK starter connects an AG-UI/LangGraph agent to team chat with native results and demonstrated approval gates for Linear and Notion writes. Its quick start uses a managed channel; operating your own runner is a separate choice. Test approver identity and tool scope in the actual workspace.
Enterprise fit: The closest channel reference for the proposed workflow: documented write interrupts and resumable approval cards. Implement eligible voters, quorum, veto, expiry and atomic execution separately; verify managed delivery and approver identity.
1,000-10,000 stars; repository created June 2026; published GitHub release.
Reviewed 5 October 2026. Star bands measure attention, not installations or enterprise reliability. Every route still needs the tenant control plane and delegated decision policy.
These projects can inform a bounded pilot or interface design. Their documented boundaries and short public histories do not establish production readiness for this operating plan.
Adapt pages, specialist roles and reviewed learning workflows.
Combine a page workspace with text, calls, Slack and scheduled work. The README describes a single-owner starting point; shared editing and automatic multi-Dot delegation remain further work. Slack, spoken compute delegation and cloud Learning delivery still need connected-service verification.
Enterprise fit: Useful page, specialist and learning interfaces. Its single-owner identity model needs replacement for multi-tenant use; specialist roles and separate conversations do not establish separate customer authority.
1,000-10,000 stars; repository created September 2026; no published GitHub release.
Adapt plans, browser takeover, receipts and recovery UI.
An alpha personal-agent app for phone and web with task plans, browser takeover, approvals, receipts and recurring checks. Its server and workers are useful recovery references. Rich Threads, live model reasoning and Google connections need separate configuration; sample flows do not verify enterprise isolation.
Enterprise fit: Useful plans, takeover, receipts and recovery UI. Adapt those surfaces to your control plane; the personal-agent design and configured sample flows do not prove multi-tenant authorization or the delegated ballot.
1,000-10,000 stars; repository created September 2026; no published GitHub release.
Inspect persistent tasks and event-watching loops in a bounded pilot.
AGPL-3.0, persistent tasks and event-watching loops; hosted or self-hosted. Assess device authority, provider data paths and the multi-service operating cost. Task persistence does not establish the proposed voting policy.
Enterprise fit: Inspect its persistent-task and event-loop design in a bounded pilot. Treat tenant enforcement, delegated approvals and recovery guarantees as acceptance tests; the public history is too short to infer operating maturity.
100-1,000 stars; repository created September 2026; no published GitHub release.
Explore local persistent agents for a trusted operator.
MIT, local persistent agents on the Pi harness. A trusted-user workspace with host-level file and shell permissions, without per-tool approval. Add isolation and external policy enforcement.
Enterprise fit: Useful for a trusted operator on an isolated host. Host-level tools and unsandboxed patch paths make it a poor direct foundation for mixed-trust tenants; external authorization and isolation are substantial work.
Under 100 stars; repository created September 2026; no published GitHub release.
Explore interface and connector patterns.
MIT, inspectable workspace with approvals and connectors. Maintainers label it a prototype: multi-user hosting and hostile-web isolation are not production ready.
Enterprise fit: Use for interface and connector exploration. The documented multi-user and hostile-web limitations exclude it from the production shortlist until those boundaries are redesigned and tested.
1,000-10,000 stars; repository created May 2023; no published GitHub release.
Repository creation dates describe the public repository, not necessarily the current product's age. A missing GitHub release does not mean there are no tags or deployments. Release labels and alpha/prototype descriptions come from maintainers; enterprise-fit judgments are editorial. This review does not establish independent production outcomes or completion of the control requirements.
Runtime documentation checked 1 October 2026: OpenClaw security, Jelly, Comma and Open Dots. Product documentation is evidence of described capabilities and limitations, not independent assurance.
CopilotKit project documentation checked 4 October 2026. These are vendor-maintained templates and described controls, not independent production assurance. CopilotKit's OpenDots is separate from the Open Dots prototype by Anil-matcha.
Define the boundary before deploying. AG-UI connects the agent and interface; it does not grant execution authority. Self-hosting an app does not establish that Threads, Learning, model inference, voice or channel delivery stay inside your infrastructure. Map each configured service, cost and data path. A saved conversation is not a durable task ledger, and an approval card does not implement the one-hour delegated decision policy. Keep tenant isolation, default denial, expiry, explicit rejection and restart behavior in enforceable application policy.
Treat learning as a reviewed change. OpenDots documents conversation evidence routing into Learning containers and reviewed publication of skills. Version and test a proposed skill before granting it to workers; keep learned instructions separate from policy and credentials. Test deletion and retention across conversations, memory, skills and audit records using the personal-data guide.
Choose one recurring workflow and compare accepted outcomes with its human baseline. A busy agent, token savings or a faster draft alone does not establish ROI.
Measure baseline volume, completion time, quality and cost. Run in shadow mode; agree on acceptance, incident and cost thresholds before seeing results.
Pilot bounded actions with human review. Count accepted work, review minutes, approval wait, rework, failures and all operating costs. Stop or redesign if quality or controls fail.
Scale only after stable quality, tested recovery and positive net value. Recheck each new tenant and workflow. Review avoided costs separately from capacity that has not become cash savings.
Full-cost ROI = (realized benefit - total cost) / total cost. Total cost includes setup, models, infrastructure, integration, human review, security, maintenance and incident recovery. Track cost per accepted task alongside quality and escalation rate.
Illustration, not measured evidence: 400 accepted tasks/month saving 15 minutes at €50/hour yield €5,000/month of capacity value. If monthly review costs €1,500 and all other recurring costs €1,000, net capacity value is €2,500/month. A €7,500 setup pays back in three months if that capacity is realized. Over six months: €30,000 benefit, €22,500 total cost, €7,500 net benefit and 33% full-cost ROI. Unused capacity lowers the realized benefit.