Running Self-Hosted AI Agents
Self-hosted AI agents keep models, data and execution in your environment. Learn the infrastructure, security, costs and operating trade-offs before deployment.
Self-hosted AI agents run the model, agent runtime, data connections, and task execution inside infrastructure you control. That may mean your own cloud account, private network, on-premises servers, or a mix of all three. Teams choose this route when sensitive data, regulatory obligations, predictable usage costs, or integration control matter more than the speed of a hosted service.
The trade-off is simple. You gain control, but you become responsible for uptime, security patching, model serving, access controls, monitoring, and incident response. For most teams, a self-hosted agent is not a weekend project. It is an operating capability.
If you are still deciding whether private infrastructure is justified, our guide on self-hosting versus API use covers the practical trade-offs teams run into after the initial proof of concept.
What makes an agent self-hosted?
A chatbot becomes an agent when it can take actions across systems, follow a multi-step process, use tools, remember relevant context, and hand work back to a person when needed.
A self-hosted agent does those things within an environment you administer.
That can include:
- A model server that generates responses
- An agent runtime that plans steps and calls approved tools
- A retrieval layer that searches internal documents or business data
- Connectors to systems such as ticketing, CRM, email, databases, and file storage
- A policy layer that restricts what actions an agent can take
- Logging, monitoring, identity management, and audit records
- A human approval process for actions with financial, legal, customer, or security impact
The phrase “self-hosted” is often used too loosely. It does not always mean every component remains in your environment.
For example, you might host the agent workflow and your data retrieval layer privately, while sending prompts to a third-party model API. That is a private agent architecture, but not a fully self-hosted stack. Or you might run the model privately but use a hosted observability service. Neither approach is automatically wrong. The important point is knowing where data travels, where it is retained, and which vendor has operational access.
Before adopting any platform, ask for a clear architecture diagram. It should show the model endpoint, databases, tool calls, logs, telemetry, backups, support access, and outbound network paths.
Options for running self-hosted agents
There are three common approaches. The right one depends on your risk profile, technical capability, and how much control you genuinely need.
1. A packaged self-hosted agent platform
Platforms in this category provide a user interface, workflow builder, agent runtime, integrations, permissions, and deployment tooling. Products such as Zanus may appear on shortlists for teams that want a more packaged route rather than assembling every system themselves.
The question is not whether a platform says it supports self-hosting. The question is what you are actually hosting.
When comparing Zanus or any similar option, confirm:
- Whether the model can run in your chosen environment
- Whether prompts, outputs, traces, and user data ever leave your network
- Whether you can disable external telemetry
- How secrets are stored and rotated
- Whether the platform supports your identity provider and role-based access controls
- Whether tool permissions can be restricted by team, workflow, and environment
- How updates are installed, tested, and rolled back
- Whether support personnel can access your instance
- What happens if you end the contract
A packaged platform reduces engineering work. It can also create a dependency on that platform’s release cycle, connector design, and security model.
This approach fits teams that have clear use cases and want to move faster without building an agent control plane from scratch.
2. A composable private stack
A composable approach uses separate components for model serving, retrieval, orchestration, storage, authentication, and observability. Your team chooses each part and connects them.
This gives you more flexibility. You can change the model layer without rebuilding workflows. You can isolate the data retrieval service from action-taking tools. You can create separate environments for development, testing, and production.
It also creates integration work. Someone must own dependency updates, schema changes, access policies, performance testing, and service reliability.
This is usually the right route when agents will become part of a core operating process, such as:
- Reviewing inbound security alerts before escalation
- Producing account research from approved internal and external sources
- Drafting responses to support cases for human review
- Collecting operational data across systems and creating incident summaries
- Processing documents that cannot be sent to a public endpoint
A composable stack should not mean every team builds everything from zero. It means you keep boundaries clear and avoid placing every responsibility inside one opaque tool.
3. A hybrid agent architecture
A hybrid setup keeps sensitive context, retrieval, workflows, and tools within your environment while using an external model endpoint for certain reasoning tasks.
This can be a sensible first step. You retain control over the highest-risk data paths without taking on the full burden of operating model infrastructure.
It is especially useful when the workflow needs access to private systems, but the underlying task does not require sending raw customer records, financial data, source code, or confidential documents outside your boundary.
The design challenge is data minimisation. Do not simply pass the entire ticket, customer record, document library, or database result to the model. Extract the smallest approved set of information required for the step.
The infrastructure a self-hosted agent needs
Teams often focus on the model and forget the surrounding operating environment. The model is only one service in the system.
A production agent generally needs the following layers.
Compute and model serving
You need enough compute capacity for the response time, concurrent users, context size, and reliability target your workflow requires.
A document classification agent that works asynchronously has very different needs from a voice agent that must respond in a live conversation. Start with the workflow. Define what response time is acceptable, how many tasks can arrive at once, and what should happen when capacity is unavailable.
Plan for:
- Production capacity and headroom for peak demand
- A separate environment for testing changes
- Health checks and automatic restart procedures
- Queues for non-urgent jobs
- Rate limits to prevent a single workflow from consuming all capacity
- A fallback process when the model service is down
Do not put an agent into a customer-facing workflow until you have tested degraded conditions. A slow or unavailable model should lead to a clear fallback, not a partially completed action.
Data retrieval and storage
Agents need context, but unrestricted access to internal information is a security problem.
Use retrieval boundaries that match existing permissions. If a staff member cannot view a contract, the agent should not retrieve it on their behalf. This sounds obvious, yet many early implementations index broad internal folders and apply access checks only at the front-end interface.
The access decision needs to happen when documents are retrieved, not only when a user opens the final response.
You also need retention rules for:
- Prompts and outputs
- Conversation histories
- Tool-call logs
- Retrieved documents and snippets
- User feedback
- Evaluation data
- System backups
Treat agent traces as sensitive operational data. A trace can contain customer information, account numbers, internal policies, or credentials accidentally exposed by a connected system.
Identity, secrets, and permissions
Every agent should have a distinct machine identity. Avoid one shared administrator credential that gives every workflow broad access to every system.
Use separate permissions for:
- Reading data
- Writing records
- Sending messages
- Making financial or contractual changes
- Accessing production systems
- Managing agent configuration
Secrets should be stored outside prompts, workflow files, and source repositories. They should be scoped to the smallest possible permission set and rotated on a defined schedule.
If an agent can call a tool that sends email, creates tickets, changes records, or triggers deployment processes, that tool needs a policy boundary. The model should propose an action. A controlled service should decide whether that action is permitted.
Observability and audit records
When an agent does something unexpected, you need to reconstruct why.
Keep records of the user request, the policy applied, the data sources retrieved, tools proposed, tools actually called, outputs produced, and approvals granted or denied. Redact sensitive content where possible, but preserve enough detail for investigation.
A useful operational dashboard tracks:
- Failed tool calls
- Timeouts and queue depth
- Approval rates
- Escalations to people
- Policy denials
- Repeated failed attempts
- Workflow completion rates
- Cost by workflow or business unit
The most important metric is not how many tasks the agent started. It is whether it completed approved work accurately and safely.
Where self-hosted agents go wrong
The biggest failures are usually not model failures. They are design and operating failures.
Too much authority too early
A common mistake is giving an agent permission to act before it has proved it can make sound recommendations.
Start with read-only workflows. Then move to draft outputs. Then allow low-risk actions with clear guardrails. Only after sustained testing should you consider higher-impact actions.
For example, an agent can summarise a support case, identify likely knowledge base articles, and draft a reply. A person reviews the draft before it is sent. That produces useful time savings without putting customer communication on autopilot.
Prompt injection through business data
An agent may read content from an email, uploaded document, website, or ticket that contains instructions intended to manipulate it. The content might say to ignore policies, reveal confidential data, or call an unrelated tool.
Do not treat retrieved content as trusted instructions. Separate system rules from retrieved data. Limit the tools available to each workflow. Validate tool inputs before execution. Put approval steps around consequential actions.
No defined owner
Self-hosted infrastructure needs a named operational owner. Without one, security updates are delayed, broken connectors remain broken, logs are not reviewed, and cost grows unnoticed.
Ownership is not one person doing everything. It means someone is accountable for the service level, change process, risk register, and escalation path.
Building before proving the workflow
Do not begin with a general-purpose agent that can access every internal system. Begin with one narrow business process that has a clear input, output, owner, and success measure.
Good first workflows tend to be repetitive, bounded, and easy to review. They have known sources of truth and a clear human escalation path.
Our Enterprise DNA learning resources can help teams build the operational skills around workflow design, evaluation, and governance before expanding access.
What does it cost to run?
The cost of a self-hosted agent is more than compute.
Your operating cost has five main parts:
- Infrastructure for compute, storage, databases, networking, backups, and disaster recovery.
- Engineering time for deployment, integration, testing, upgrades, and incident response.
- Security and compliance work for identity, logging, reviews, access controls, and vendor assessment.
- Operations for monitoring quality, correcting failures, managing data, and supporting users.
- Change management for training staff, documenting new processes, and maintaining human review.
A low-volume internal workflow can be inexpensive in raw infrastructure terms and still be costly if it requires frequent engineering attention. Conversely, a high-volume workflow can justify dedicated capacity if it replaces a slow, repeatable manual process.
Build a cost model before deployment. Include expected task volume, average processing time, concurrent demand, staff review time, storage retention, support effort, and a contingency for failures. Compare that total with your current process cost, not only with a third-party API bill.
A practical deployment sequence
A sensible rollout is staged.
- Choose one workflow. Define the trigger, source systems, expected output, approval point, and failure path.
- Classify the data. Identify what the agent can read, what must stay private, and what must never enter prompts or logs.
- Set permissions first. Create dedicated identities and least-privilege tool access before connecting systems.
- Run in shadow mode. Let the agent produce recommendations while people continue the existing process.
- Evaluate real cases. Test routine cases, edge cases, malicious inputs, missing data, and unavailable systems.
- Require approval for actions. Keep a person in the loop for external messages, record changes, or financial impact.
- Release in small groups. Start with a limited user group and monitor failures closely.
- Review weekly. Look at errors, cost, policy denials, user feedback, and cases requiring escalation.
- Expand authority slowly. Only automate an action after evidence shows the workflow is reliable and controls work as intended.
For teams that need to operationalise agents across several workflows, Omni provides a broader view of how agent-enabled processes can fit into the business. The Omni Ops approach is particularly relevant when the challenge is running, monitoring, and improving those processes over time.
Self-hosting is worth the effort when control of data, permissions, and operating behaviour is central to the use case. It is not automatically better because it is private. The right choice is the one your team can secure, maintain, evaluate, and support after the pilot excitement has passed.
If you want to assess whether a private agent architecture fits your operating model, book a call with Sam.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideEDNA Learn
Start free on EDNA Learn
Free account, no card. Run the Claude Code and agent-building course and start earning MENTOR credits.
Start freeFree Resource
Get the AI Operating Layer guide
Connect the tools you already run your business on, let agents handle the repeatable work, and keep a person on the gate.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the Guide