Self-Hosted AIOps Agent: Watches Your Systems, Finds What Broke

Custom AIOps agent · deployed on your infrastructure

Ask your whole stack anything. Or let it watch, and tell you first.

One agent, deployed in your environment, wired into everything you already run: logs, metrics, events, system health, internal and external APIs, databases, and your code. It reads your systems and answers in plain language. It never changes them.

Fixed price, no hourly surprises. Delivered in 60 days, usually sooner. Paid in four milestones, you release each one only when you approve it.

It is 3am and something is on fire. You have fifteen Grafana tabs open, you are grepping logs by hand, and you still cannot say what actually changed. Dashboards show you data. They do not tell you what it means, and they never wake you up with the answer.

How it works

One agent. Every source. Plain language.

Two ways to use it, one build, all wired for you. I do 100% of the wiring: every connector, every query, every mapping. What I need from you is read-only access to each system, and nothing else.

01

It connects to everything

Logs, metrics, events, system health, internal and external APIs, databases, and your code. I wire all of it; you hand over one read-only credential per system.

02

You ask, or it watches

Chat with it in plain language from a web UI, your own app, or your MCP client. Or let it watch your systems live and decide for itself what is worth a look.

03

It hands you the answer

Not “CPU is high” but why, since when, what changed, and what to do, plus the one chart that proves it, routed to wherever your team already lives.

AI model — bring your own hosted API — frontier or open‑weight or self‑host an open model for full privacy YOUR ENVIRONMENT runs in your own infrastructure The Agent one binary — tooling built in built‑in data + compute tooling connects · queries · combines · computes agentic — plans · correlates · hunts deeper exposes: OpenAI API + Agent MCPchat on demand · autonomous on triggers no separate services to wire up Your data sources Logslog files · Loki · traces MetricsPrometheus · counters · KPIs Eventsapp + business event streams System metricsCPU · disk · memory · load Internal APIsREST · gRPC · services Coderepos · trace → exact line ① PULL — ask on demand Your MCP client Claude Code · Cursor via Agent MCP — one “chat” tool Web chat (Chatz) your team · plain chat ② PUSH — autonomous Triggers schedules + event hooks error spikes · disk full · cron fires itself — no human in the loop Alerts out Slack · Discord · Telegram · email reports out beyond your env the model reasons &calls the agent's toolshosted → only calls leave pull push report → any channel
One agent, deployed in your environment. It reads your sources, reasons across them, and answers on demand or acts on its own.

What you get

One complete build. Everything included.

No tiers, no locked features, no upsell to reach the good part. The full agent, wired to your stack and deployed on your own infrastructure.

Ask

Answers, not dashboards

  • Up to 20 data sources wired for you: APIs, databases, logs, metrics and events, queues and streams, cloud provider APIs, object storage, and code repositories
  • Cross-source reasoning that correlates across all of them in one pass, so a question spanning your database, your logs and last week’s deploy comes back as one answer
  • Server-side computed metrics, real joins, rates and aggregations, not sampled
  • Three ways in, a web UI that builds charts fresh per answer, an OpenAI-compatible API, and an Agent MCP
  • Unlimited team seats, no per-seat billing
Watch

It comes to you first

  • Autonomous investigation that traces a symptom back to the commit or setting that caused it, and writes out the exact fix
  • 10 event-driven hooks, tuned to what abnormal means on your stack
  • 20 scheduled digests, on whatever cadence you want
  • 10 alerts routed by severity to Slack, Discord, Telegram, email or any channel with a webhook or API
  • 5 custom playbooks for the failure modes you already know by heart
Own

Yours, forever

  • A repository with a Dockerfile, shipped through your normal review and your normal pipeline
  • Open source building blocks, read every line, audit it, fork it
  • Runs on your infrastructure, your data stays on your own hardware
  • Bring your own model, anything speaking an OpenAI-compatible or Anthropic-compatible API, hosted or self-hosted; you pay your provider directly, I take no cut
  • Deployment and hardening done for you, plus a tuning window after handover

It ships through the same reusable CI (GitHub Actions) as everything else I publish.

Not billed separately. Every build is checked before it reaches you, so you can verify exactly what you are running.

Code checks

Linting, tests and vulnerability scanning on the code, every commit.

Image scanning

A Grype vulnerability scan on the container image before it ships.

Supply-chain proof

An SBOM and provenance attestation on every image, so nothing about the build is a mystery.

It reads your systems. It never changes them.

Every credential the agent uses is read-only. Where it reads a database it uses a read-only user you create, or every query runs in a read-only transaction. It investigates, traces the cause, and hands you the fix as a diff with the reasoning. It never applies that change: no commits, no pull requests, no config edits. You decide what ships. Nothing leaves your network except the LLM call you configure; self-host the model and nothing leaves at all.
$15,000one complete build · delivered in 60 days, usually sooner

No tiers, no locked features, no seats, no subscription. You are paying for the engineering: senior integration work, proven open-source blocks assembled onto your exact stack, the bespoke connectors, the tuning, and a working handover. Bigger estate? Extra sources, hooks or playbooks are flat per-unit add-ons, quoted up front, never surprise scope.

  1. 1

    Scoping and access $2,500

    Written scope, confirmed source list, credentials in place, repository created with CI running.

  2. 2

    Core integration $5,000

    First batch of sources wired in; web UI, API and MCP running in your environment; first working demo.

  3. 3

    Full wiring and watch layer $5,000

    All sources in scope; hooks, digests, alerts and playbooks configured; container built and deployed.

  4. 4

    Handover and tuning $2,500

    Tuning window against real traffic, documentation, handover.

Paid in four milestones. Each one is agreed before it starts and paid only when you approve its deliverable, so you never pay up front for work you have not seen.

Why not just more dashboards

Dashboards show data. This understands it.

It gives the answer

It correlates across sources and hands you the conclusion, plus the one chart that proves it, instead of a wall of panels you still have to read yourself.

It is proactive

It watches and comes to you before you think to look. Not one more screen you have to remember to check.

Built for your stack

No forcing your data into someone else’s schema. It reads what you already run, and plugs into your Grafana and Datadog rather than replacing them.

You own it. No rent, no leash.

🔓

Open source, audit everything

The reusable building blocks are open under a permissive license on my GitHub, audit them yourself. Your custom integration on top ships to you as full source. Nothing you run is a black box.

🏠

Runs on your infrastructure

Agent, MCP, connectors and UI all live on your hardware. The only thing that leaves is the LLM call you configure. Self-host the model and nothing leaves at all.

🛡

Paid in phases, never up front

Four milestones, each agreed before it starts and paid only when you approve its deliverable. You never pay up front for work you have not seen.

Three ways to engage, see below ↓

∞

No seats, no subscription

Pay once to build it, then run it forever. Agent Care, the monthly maintenance retainer, is there if you want it. A choice, never a requirement.

After handover, if you want it

Agent Care: keep it healthy without lifting a finger.

The build is yours to run forever with no dependency on me. But if you would rather not own the upkeep, Agent Care is an optional monthly retainer that keeps the agent current, connected, and quiet. $1,000/month, no minimum term, cancel any month.

Updates and security patches

Core updates as a pull request or a tagged image, plus dependency bumps, CVE fixes and base-image rebuilds, each shipped through the same CI pipeline as the original build: linting, tests, Grype scan, SBOM and provenance.

Connector repair

APIs change, schemas drift, log formats shift, keys rotate. I fix the connectors before any of that turns into a gap in what the agent can see.

Health monitoring and tuning

I watch the agent itself, logs, crash loops, rate limits, token burn, stuck queues, and tune thresholds, hooks, alerts and prompts so false positives stay down and the signal stays worth trusting.

It includes a direct channel with a five-hour response time and up to six hours of change work a month. It does not cover incident response (the agent reports, your team acts), application debugging, or brand-new sources, hooks or playbooks, which are quoted as add-ons. Agent Care runs the same three ways as the build, see below.

Straight answers

What does the price cover?

One complete build: up to 20 data sources wired in, the watch layer with its hooks, digests, alerts and playbooks, hardening, full source as a repository with a Dockerfile, deployment by your team or by me, and a tuning window after handover. Extra sources, hooks or playbooks beyond the included counts are flat per-unit add-ons, quoted up front.

How does payment work?

In four phases: scoping and access ($2,500), core integration ($5,000), full wiring and watch layer ($5,000), handover and tuning ($2,500). A phase is agreed before it starts and paid when you approve that phase’s deliverable, so you are never paying up front for work you have not seen. That shape holds whichever way we engage, whether the money runs through your own contract and invoicing or through a platform’s milestone escrow.

How do we actually engage?

Three ways, and I do not care which one you pick. Direct B2B on your own contract and invoicing, which is usually simplest if your company already has a procurement process. Or through Contra or Upwork if you would rather have the platform hold the milestones and handle the paperwork. Same work, same phases, same price.

What do you need from my team?

One scoping call to go through what you run and what you want watched. Then one read-only credential per system, created by you: a database user, a scoped API key, a read-only cloud role, whatever your policy allows. After that I do not need anything from you until handover, apart from the occasional question about what a metric or a table means in your business. For running it during the build, your team ships the container through your normal pipeline, or you hand me a deploy target and I do it, your call. And if your policy wants an NDA signed before we talk specifics, that is fine.

What counts as a data source?

Anything with a machine-readable interface I can read using a read-only credential: a documented API (REST, GraphQL or gRPC), a database, a log, metrics or event system with a query API or exporter, a message queue or stream, a cloud provider’s API, object storage or a file share with structured files, or a git repository. Internal systems count as long as they expose one of those, and a custom connector for an internal API or database is part of the build. What does not count: anything reachable only through its user interface. Legacy applications with no API and no database access, desktop-only clients, and anything that would need browser automation or scraping are outside the standard scope. If you have one, we look at it during scoping and either find a real interface for it, quote it separately, or leave it out.

Is my data safe? Can it change or break anything?

The agent, the MCP, the connectors and the UI all run on your own infrastructure. It is read-only by design: it never writes to your systems, it investigates and hands you the fix rather than applying it, and it has no write access to your files, data or config. Where it reads a database it uses a read-only user you create, or each query runs in a read-only transaction. The only thing that leaves your network is the LLM call you configure, and self-hosting the model means nothing leaves at all.

Can it be wrong?

Yes, it reasons with an LLM, and an LLM can misread a situation. The system is built around that fact: every conclusion ships with its evidence, the chart, the trace, the query results, so you verify in seconds instead of taking its word. It never acts on its own conclusions, a wrong answer costs you a glance, not an outage. And the tuning window exists precisely to adjust hooks and thresholds against your real traffic until the signal is worth trusting, noisy alerts get tuned out, not shrugged at.

Which model does it use? Does my data go to a third party?

Your choice, as long as it speaks one of two API styles: OpenAI-compatible or Anthropic-compatible. Those two shapes are what the industry standardised on, so it covers the big providers, the aggregators and the newer players (many of which serve both styles), plus anything you run yourself behind an OpenAI-compatible endpoint, which is what vLLM, Ollama and similar servers expose. The only thing that leaves your network is the LLM call you configure; self-host the model and nothing leaves at all. You pay your provider directly at their rates, and I take no cut of it.

What does it cost to run?

The only meaningful running cost is the model. Token spend depends on which model you pick, how much the watch layer investigates, and how hard your questions are, so I estimate it for your actual setup during scoping and tune hooks and digests so the burn stays proportional to what you get out of it. Self-host the model and the token bill disappears entirely; the rest is a container on hardware you already run.

What happens after handover?

A tuning window is included, where I adjust hooks, thresholds and playbooks against your real traffic. After that it is yours: full source, running on your hardware, no dependency on me. After that, if you would rather not own the upkeep, Agent Care picks it up on a monthly retainer: updates and security patches, connector repair when an upstream API changes, health monitoring, and continued tuning. Never required.

Do I own it, or am I locked in?

You own it. The building blocks are open source under a permissive license, you get full source for your integration, connectors, playbooks and config, and it all runs on your hardware. Pay once to build it, then run it forever. No seats, no subscription, no rent. Agent Care exists if you would rather not maintain it yourself, and it is always a choice, never a requirement. If I vanish tomorrow, your system keeps running.

Tell me about your stack.

This is for teams running enough infrastructure that something can break and nobody notices for hours. Send me what you run and what “watch my business” means for you. We get on a call, I scope it, and I build it into your stack. You end up with a repository and a container your team can deploy anywhere, or I deploy it for you.

Get in touch →

Three ways to run it, whichever suits your procurement: direct B2B on your own contract and invoicing, reach me on LinkedIn and we sort the terms directly. Or through a platform that holds the milestones for you, on Contra or Upwork. Same work, same phases, same price.