AI coworkers are here. Where custom AI agents still have to take over
Off-the-shelf agents now handle email, scheduling, research and popular apps well. Your own systems are where they stop, and where custom agent engineering starts.
Off-the-shelf agents like OpenAI’s dots and Meta’s Muse now handle generic work in popular apps well; custom AI agents take over where the work runs through your own systems, data and rules.
A custom AI agent is software that uses a language model to plan and take the next step in one of your workflows, with access, permissions and checks engineered for your systems. On September 29, 2026, OpenAI launched dots, which its dots launch post describes as “always-on agents”, and Meta launched Muse for Small Business. Both plug into the software most companies already share: Meta names Slack, Shopify, QuickBooks and Notion among Muse’s connectors, and OpenAI says dots can connect to over 4,000 apps. Neither arrives knowing your custom CRM, your legacy ERP, the data scattered across your internal tools or the compliance rules your business runs on.
That is where generic agents stop being useful, and OpenAI draws the same line in its own launch post. Its specialist dots, the ones with deep integrations into a company’s systems of record, start as “focused enterprise pilots”, with OpenAI’s engineering teams working directly with each organization to define what the dot is responsible for, which tools it can use and how people approve its work.
For a CTO or founder, the practical question after the launches is which workflows to hand to these agents now and which ones need engineering first. The answer depends less on the agent than on where the work lives.
What did OpenAI and Meta launch?
On September 29, 2026, OpenAI introduced dots, always-on agents that live in ChatGPT and work from their own cloud computer, together with ChatGPT Space, a shared space where teammates, ChatGPT and their dots build on the same knowledge. The same day, Meta launched Muse for Small Business, a set of new skills and connectors for Muse, the personal AI agent it introduced on September 8. Both companies are selling the same idea: an agent you hand a goal to, which works across the apps you already use and checks with you before anything consequential.
OpenAI dots
OpenAI says dots are powered by GPT-6 Astra, have their own cloud computer and browser, and “can readily connect to over 4,000 apps” through its plugin ecosystem. Each person starts with one primary dot, and OpenAI is previewing specialist dots that take on dedicated responsibilities inside an organization. You can reach your dot in ChatGPT, Slack or Microsoft Teams.
The first dot is included in the Pro and Business Premium plans at no extra cost. At launch, Pro users in the European Economic Area, Switzerland and the UK can’t use dots, Business Premium users can in every supported ChatGPT region, and Enterprise workspaces get a beta that is off by default, according to OpenAI’s getting-started guide for dots.
Meta Muse for Small Business
Muse runs on what Meta calls Muse Secure VM, a dedicated virtual machine with its own browser. Muse for Small Business adds connectors for Asana, Box, Canva, Dropbox, Figma, Granola, HighLevel, Intuit QuickBooks, Klaviyo, Lovable, Notion, Shopify, Slack, Stripe and Zoom, plus Facebook and Instagram business accounts, and Meta says Muse “also supports custom connectors so you can plug in services we don’t support yet”. Muse is free for most of what people need, with subscription plans for more. Meta’s Muse announcement says the agent checks with people before sensitive actions and shows a complete audit trail of everything it has done and plans to do. It is available in the US and Canada.
The enterprise platforms behind them
Both launches sit on larger enterprise pushes. In February 2026, OpenAI introduced Frontier, a platform for enterprises to build, deploy and manage agents, which it says helps teams move “beyond isolated use cases to AI coworkers that work across the business”. On September 28, Meta launched Meta Enterprise Platform to bring the Muse agent, Meta Business Agent, the Muse API and Muse Code to businesses and developers. Meta said in June that more than one million businesses already use Meta Business Agent on WhatsApp and Messenger to respond to customers around the clock.
Here is how the two launches compare on the points that matter for this decision:
| OpenAI dots | Meta Muse for Small Business | |
|---|---|---|
| What it is | An always-on agent in ChatGPT with its own cloud computer and browser | New skills and connectors for Muse, Meta’s personal AI agent |
| What it connects to | Over 4,000 apps through plugins, reachable in Slack and Teams | Dozens of business tools, 15 named at launch, plus Facebook and Instagram accounts and custom connectors |
| Who can use it | Pro (outside the EEA, Switzerland and UK), Business Premium, Enterprise beta | People in the US and Canada |
| Pricing model | First dot included in Pro and Business Premium | Free for most use, subscription plans for more |
| Control model | Built-in approval rules, Custom Rules and auto-review | Approval before anything publishes, sends or spends, plus an audit trail |
| Your systems of record | Specialist dots, starting as focused enterprise pilots | Custom connectors that someone has to build |
What off-the-shelf agents already do well
Off-the-shelf agents are now good at general computer work: browsing, filling in forms, researching, drafting and moving information between popular apps. Stanford HAI’s 2026 AI Index reports that AI agents went from 12% to about 66% task success on OSWorld, a benchmark of real computer tasks across operating systems. On WebArena, a set of long web tasks, success reached 74.3% in early 2026, within 4 percentage points of the human baseline of 78.2%, according to the report’s technical performance chapter.
For a small company, that covers a lot of real work, and the launches make it cheap to try. An agent can triage an inbox and draft replies, summarize a week of Slack threads, research a prospect before a call, turn meeting notes into tasks in Asana or Notion, or prepare a draft invoice for someone to approve. OpenAI’s own launch examples include a dot that noticed an early tester had forgotten to invoice a publication, prepared the invoice and sent it after his approval. Each of those jobs runs inside software the vendor has already connected, on data that software already exposes, behind an approval step the vendor has already designed.
Our advice to most small companies is to hand that work to an off-the-shelf agent now. The cost sits inside a subscription you may already pay for, the vendor maintains the integrations, and the approval models are sensible for low-stakes actions. Meta says nothing publishes, sends or spends without your approval in Muse. Dots start with built-in rules for when to act on their own and when to ask, and OpenAI’s Custom Rules let you allow specific actions, require approval for them or block them. Keep one line from OpenAI’s launch post in view all the same: “Dots can still make mistakes, so always review consequential work.”
Where do generic agents stop?
Generic agents stop at the edge of what their vendor has connected, can see and can check. A custom CRM, a legacy ERP, internal data and compliance rules all sit past that edge, which is why the hard part of putting an agent to work in your business is access and context rather than the model. Anthropic reached the same conclusion from usage data in its September 2025 Economic Index report: deploying AI for complex tasks “might be constrained more by access to information than on underlying model capabilities”.
Companies with full IT departments report the same wall. In Salesforce’s 2026 Connectivity Benchmark, a survey of 1,050 IT leaders at organizations with at least 1,000 employees, 96% agreed that AI agent success depends on seamless data integration across all systems, yet only 27% of their applications were integrated with each other. The top hurdles they named were risk management, compliance, security or legal implications (42%), a lack of internal expertise in AI and agent design (41%), legacy infrastructure or system incompatibility (37%) and integrating siloed apps and data (35%). A 50-person company faces the same four problems with fewer people to solve them.
The new agents’ own documentation shows where the edge sits. OpenAI’s admin guide for dots states: “Enabling dots does not grant access to every app or website.” It also notes: “A dot’s cloud computer does not automatically inherit a member’s local VPN, browser sign-ins, or device policies.” If your ERP sits behind a VPN, a dot can’t reach it until someone builds a way in.
Benchmarks point the same way. The 2026 AI Index notes that agents “still fail roughly 1 in 3 attempts on structured benchmarks”, and those benchmarks are public, documented environments. Your systems add something no benchmark has: years of decisions about what each field, status and exception means in your company. Four kinds of systems account for most of the gap.
A custom CRM: the data model only you have
A custom or heavily customized CRM holds the fields, statuses and relationships that describe how your company sells to and serves its customers, and no plugin directory knows what they mean. An off-the-shelf agent can read a standard CRM through a published connector, but it has no way of knowing that your “Tier 2” flag marks a customer on a legacy contract, or that a deal in “Review” needs finance sign-off before anyone touches pricing.
Salesforce AI Research measured how hard realistic CRM work is for agents. In CRMArena-Pro, a benchmark of nineteen expert-validated sales, service and quoting tasks published in 2025, the leading LLM agents of the time achieved “only around 58% single-turn success”, with performance dropping “to approximately 35% in multi-turn settings”. The same agents showed “near-zero inherent confidentiality awareness”. Models have improved since, and the benchmark contains none of your customizations. What it shows still holds: CRM work is multi-step, conversational and full of data an agent must not leak.
Anthropic’s report uses a CRM to explain why context matters. Developing a sales strategy for a key account, it says, “might require Claude having access not only to information contained within a Customer Relationship Management system, but also to tacit knowledge located in the minds of account executives, marketers, and external contacts.” A generic agent connected to your CRM gets the first half at best.
What makes a CRM agent work is an access layer that exposes your objects with their meaning attached: which fields the agent may read, which it may write, which actions need a person, and what each status implies downstream. That layer also has to catch failures that don’t announce themselves. A write can return success while the record it was meant to change stays as it was, and an agent acting through that call has no error to react to. So the access layer should check each write against the system of record, not only the response code, and log both.
A legacy ERP: no plugin, no API, behind a VPN
A legacy ERP is the hardest place for an off-the-shelf agent to work, because it often has no plugin, no modern API and no route in from the public cloud. Plugin directories are built around popular SaaS products. An on-premises ERP installed a decade ago, or an older edition your team extended in-house, isn’t in them, and a dot’s cloud computer starts outside your network.
There are three ways to give an agent reach into a system like that, and they trade cost against fragility:
- Use an interface that already exists. Older systems often expose a web service, a database view or a scheduled file export that nobody has wired to anything yet. Wrapping it is usually the cheapest reliable route.
- Put a small integration service in front. When the interface is awkward or too permissive, a thin service turns tables, stored procedures and files into a short list of clean, permissioned actions an agent can call, and it logs every call.
- Drive the user interface as a last resort. An agent can click through screens the way a person does. That works where nothing else does, and it breaks when a screen changes.
We learned the limits of the third route on our own systems. Our internal AI web agent, a proof of concept built for our own team, takes a text prompt, navigates unicrew’s own project management system, finds the people whose time logs are missing for the previous week, and notifies them and their managers in Google Chat. It improved timely work-time logging compliance across our team by 30%. It also showed how much a browser-driven agent depends on stable UI selectors, which is why we take an API wherever one exists.
Whichever route you take, put it behind one permissioned interface instead of wiring each agent separately. That is what an MCP server does, and it is how the same ERP access can later serve a custom agent, a support assistant and any general-purpose agent your team adopts that supports custom connectors.
Internal data: context an agent can’t see
Most of what makes a decision correct inside your company lives in data no generic agent can see: shared drives, wikis, ticket histories, spreadsheets, call notes and the experience of people who have done the job for years. Off-the-shelf agents learn about you from what you connect and what you tell them, so their judgment is only as good as the slice of context they were given.
Anthropic’s Economic Index report names the cost of closing that gap. For some firms, it says, “costly data modernization and organizational investments to elicit contextual information may be a bottleneck for AI adoption.” Connecting a shared drive is easy. Knowing which version of a pricing sheet is current, which wiki pages are obsolete and which customer notes are confidential is the data engineering an agent depends on.
Context also has to be governed. When an agent learns from your data, you need to know what it kept and be able to remove it. OpenAI’s dots privacy and safety FAQ says you “currently cannot view, delete or directly modify individual dot memories, including specific details that enter the dot’s context from plugins”. The same FAQ notes that “Human review may occur in limited circumstances, including safety-related cases, even when model improvement is turned off.” Meta says Muse users can tell it to forget specific things it has learned. Either way, the memory lives in the vendor’s cloud, tied to the person who owns the agent. That needs a plan before customer records go anywhere near it, especially if a customer asks you to delete their data or a contract limits where their records may be processed.
A custom agent reads internal data through retrieval you control: which sources it searches, which documents each role may see, how long anything is kept and where it is processed. Where the data an agent needs isn’t reachable yet, data engineering comes first in the plan.
Compliance rules: enforced in code, not in a prompt
Compliance rules are where a generic agent’s approval model runs out. Off-the-shelf agents ask before consequential actions and let you add rules in plain language, but a rule like “refunds above a set amount need a manager” or “patient records never leave the EU” has to hold every time, so it belongs in the system the agent calls.
The OWASP Top 10 for LLM Applications 2025 says the same under Excessive Agency: “Implement authorization in downstream systems rather than relying on an LLM to decide if an action is allowed or not.” OWASP says the root cause is typically one or more of three: excessive functionality, excessive permissions and excessive autonomy.
OpenAI is candid about the limits of protection inside the agent. Its FAQ says the safeguards against prompt injection, meaning instructions hidden in a website, email or document, “help reduce the risk of malicious instructions causing an unwanted action, but they do not eliminate it”. Benchmarks with policy rules make the same point. On τ-bench, which tests agents against databases and policy constraints in retail and airline scenarios, leading models score between 62.9% and 70.2% on a first attempt, according to the 2026 AI Index. That may be acceptable for drafting support replies a person reviews. It isn’t acceptable for a refund policy or a data-residency rule.
The new agents take approvals seriously. Dots come with strong defaults for which actions need your approval each time, which you can approve in advance and which a dot can take without asking, and Muse keeps an audit trail of what it has done. Those controls are built around the person using the agent: in a ChatGPT workspace, “Only a dot’s owner can direct it.” Company compliance usually involves someone else, such as a manager who signs off refunds above a limit or finance approving a credit note, and it needs evidence recorded in your own systems, where an auditor will look.
In a custom agent, the rules live in the access layer the agent calls. The agent proposes a refund; the service checks whether this user, this amount and this customer allow it, routes the request to the right approver, and logs the decision where your auditors already look. The model never decides whether an action is allowed. This is also the practical answer to whether AI agents can be deterministic: the model isn’t, but everything it’s allowed to touch can be.
Custom AI agents vs off-the-shelf agents, side by side
Off-the-shelf agents win on speed, price and maintenance for work inside popular apps; custom AI agents win wherever the work depends on your own systems, data or rules. Between the two sits a middle option: no-code agent builders such as n8n or Microsoft Copilot Studio, and custom connectors that extend a generic agent without building a new one.
| Off-the-shelf agent | No-code builder or custom connector | Custom AI agent | |
|---|---|---|---|
| Examples | OpenAI dots, Meta Muse | n8n, Copilot Studio, a Muse custom connector | An agent built for your CRM, ERP or support desk |
| Time to first result | Minutes to days | Days to weeks | Weeks per workflow |
| Systems it reaches | Apps in the vendor’s directory | Systems with an API or a connector you build | Anything you can expose safely, including systems without a modern API |
| Where your rules live | Agent instructions and the vendor’s approval settings | Workflow steps you configure | Code in the access layer, enforced on every call |
| Where your data goes | The vendor’s cloud, under its terms | The builder platform and the model provider | A model host, region and retention period you choose |
| Who maintains it | The vendor | You, within the platform’s limits | You, or the partner who built it |
| Cost model | Subscription | Platform subscription plus build time | Build cost plus model and hosting spend |
| Best for | Inbox, calendar, research and popular SaaS | Simple flows across tools with good APIs | Workflows through a custom CRM, legacy ERP, internal data or regulated steps |
Most companies will run all three. The useful question is which workflows belong in which column, and the deciding factor is access more than features. Our build-or-buy guide for software makes the same argument for applications: buy for commodity work, build where the workflow is your advantage.
What custom agent engineering involves
Custom agent engineering is mostly the work around the model: choosing one job, reaching the systems it needs, encoding the rules, and proving the agent does the job before it acts alone. Model choice matters, but the access layer and the evaluation suite decide whether an agent survives production.
- Name one job and write down what done means. “Handle customer service” is too broad to measure. “Draft the reply to every delivery-delay ticket, with the carrier’s latest status, for a person to approve” is a job you can score. A general-purpose assistant is how a pilot stalls before anyone lets it act.
- Map every system the job touches. For each one, decide how the agent reaches it (an API, an integration service, or the user interface as a last resort) and the minimum permissions it needs. OWASP’s guidance is to execute actions on downstream systems “in the context of that specific user, and with the minimum privileges necessary”.
- Build the access layer once. Expose those systems through an MCP server, one permissioned interface that any MCP-capable agent can use. Anthropic open-sourced the Model Context Protocol in November 2024 as “a new standard for connecting AI assistants to the systems where data lives”, and OpenAI’s DevDay 2026 recap says it is adding support for the proposed MCP Events specification so plugins can start automations when something happens in a connected app. Our MCP explainer covers how the protocol works, and our MCP development team builds and secures the servers.
- Put the rules in code. Approval thresholds, data boundaries and separation of duties are enforced by the access layer, every decision is logged, and the agent cannot talk its way past them.
- Build the evaluation suite before the agent acts. Score the agent on real tasks from your own history, including the awkward cases, and rerun the suite whenever the model, the prompt or a connected system changes. Our guide to why most AI agents never reach production goes deeper on evals, guardrails and observability.
- Start supervised and earn autonomy. The agent drafts and a person approves. Each action it later takes alone is a decision you make on evidence, one action at a time.
- Budget for running it. Model and infrastructure spend continue after launch, and so does maintenance whenever your systems change.
A well-built access layer also makes off-the-shelf agents more useful. Meta says Muse supports custom connectors, and an MCP server built for your own agent can serve as that connector wherever a general-purpose agent supports MCP, under the permissions you set.
How to decide which agent gets which workflow
Sort each workflow by three questions: where the work lives, what a wrong action costs, and whose knowledge it needs. Work inside popular apps with low-stakes actions goes to an off-the-shelf agent now. Work that runs through your own systems, costs real money or compliance exposure when it goes wrong, or depends on knowledge only your team has is a candidate for a custom AI agent.
| Workflow | Where the work lives | Cost of a wrong action | Best fit |
|---|---|---|---|
| Triage the founder’s inbox and draft replies | Email and calendar | Low, every send approved | Off-the-shelf agent now |
| Weekly sales summary from Shopify and Stripe | Popular SaaS | Low, read-only | Off-the-shelf agent now |
| Answer order-status questions on WhatsApp | Messaging plus your own order system | Medium, customer-facing | Off-the-shelf front end with a custom connector to your order data |
| Update deal stages and renewal terms | Custom CRM | High, revenue and contracts | Custom AI agent |
| Match supplier invoices to purchase orders | Legacy ERP | High, money moves | Custom AI agent, rules in code |
| Prepare a customer’s data-deletion request | Internal data across several systems | High, regulatory | Custom AI agent, supervised |
Two more questions settle the close calls. How often does the work happen? Volume decides whether automating it pays back. How much does each case vary? Variation decides whether it needs an agent at all, or whether a script would do the job more cheaply and fail more predictably.
When a custom agent pays off
A custom agent pays off when the process is settled, the volume repays the build, the cases vary enough to need judgment, and the work runs through systems an off-the-shelf agent can’t reach. An agent pointed at an unsettled process amplifies the confusion instead of absorbing it, so settle the process first and automate it second.
Where unicrew fits: the systems generic agents can’t reach
unicrew builds custom AI agents for the workflows that run through a company’s own systems, and the access layer that lets any agent reach those systems safely. We build agents to each client’s requirements and sell none of our own, and we take no commission or margin on any third-party platform, so when dots or Muse is the right tool for a workflow, we say so.
Tural Mamedov, our CEO and co-founder, puts it this way:
“Every company can now rent the same generic agent for the price of a subscription, so the generic agent stops being an advantage. The advantage moves to whoever connects an agent, safely, to the systems their competitors don’t have. That connection is engineering work, and it’s the part we do.”
Our AI agent development work covers the whole build: scoping one job, the integrations, the evaluation suite, the guardrails and a supervised rollout. Our MCP work builds the access layer once, so every agent you adopt goes through the same permissioned door. For growing companies that want the model OpenAI pairs with its enterprise platform, forward deployed engineering puts one senior, AI-native engineer inside your team to own the work end to end. For startups and scaleups without an in-house AI team, that means one person who owns the workflow from the first conversation to release.
We hold ISO 27001:2022 and ISO 9001:2015 certifications, audited by Quay Audit UK, and design agent access under them. Most engagements start within two to four weeks, there is no minimum engagement period, and we work on time and materials, fixed price or team extension.
If the workflow you most want to hand to an agent runs through a system no plugin directory lists, tell us which one, and we’ll tell you what it would take to connect an agent to it safely.
Key takeaways
- Off-the-shelf agents are ready for generic work. OpenAI’s dots and Meta’s Muse handle email, research and popular apps well, and most small companies should use them for that now.
- They stop at your own systems. A custom CRM, a legacy ERP, internal data and compliance rules sit outside what an off-the-shelf agent’s vendor has connected, can see or can check.
- The vendors draw the same line. OpenAI starts its specialist dots for systems of record as enterprise pilots with its own engineers, and pairs Frontier customers with forward deployed engineers.
- Access is the hard part. Build one permissioned access layer, ideally an MCP server, and every MCP-capable agent you adopt can use it.
- Rules belong in code. Enforce approvals and data boundaries in the systems the agent calls, never in its prompt.
- Decide workflow by workflow. Sort by where the work lives, what a mistake costs and whose knowledge it needs, then start one custom agent, supervised.
Frequently asked questions
Custom AI agents are AI agents built for one company's workflows, with access, permissions and checks engineered for its own systems. Like off-the-shelf agents such as OpenAI's dots or Meta's Muse, they use a language model to plan and take steps. The difference is reach: a custom agent works inside your CRM, ERP and internal data under rules your systems enforce, while a generic agent works through the apps its vendor has already connected.
OpenAI describes GPTs as versions of ChatGPT configured for a specific purpose, while an AI agent takes actions in software to finish a task. OpenAI now calls its workspace agents an evolution of GPTs and says it plans to retire custom GPTs, recommending plugins instead. A custom AI agent goes further than either: it acts inside your own systems, under permissions and rules you define and enforce in code.
Four things set the cost: how many systems the agent must reach and how clean that access is, how consequential its actions are (which sets the evaluation, guardrail and approval work), how much data preparation the job needs, and whether someone runs it after launch. Budget for model and hosting spend after go-live too. We quote against your scope, on time and materials, fixed price or team extension. More on AI agent development.
Most of the timeline goes to agreeing the single job and reaching the systems it runs against, not to the agent itself. One job on one system with clean access moves fastest; several systems with significant data preparation take longer. A supervised pilot comes first, and autonomy grows as evaluation results justify it. Most of our engagements start within two to four weeks, with no minimum engagement period.
Use an API wherever one exists. Where it doesn't, a small integration service can turn database tables, files or older interfaces into clean, permissioned actions, and driving the user interface is the last resort because it breaks when screens change. Expose all of it through one MCP server, scope every action to the user the agent acts for, and log each call. More on MCP development.
The model inside an agent is not deterministic: the same request can lead to different steps. What can be deterministic is everything the agent is allowed to touch. Put approval thresholds, data boundaries and permission checks in the systems the agent calls, so each rule holds whatever the model proposes, and log every decision. OWASP recommends the same: enforce authorization in downstream systems rather than letting the model decide.
It depends on where your work lives. At launch, dots connect to over 4,000 apps through ChatGPT plugins and come with Pro and Business Premium plans, though Pro users in the EEA, Switzerland and the UK can't use them yet. Muse for Small Business connects to tools such as Shopify, QuickBooks and Slack plus Meta's business accounts, is free for most uses, and is available in the US and Canada. Neither reaches a custom CRM or legacy ERP without connector work.



