The Problem: Homelabs Aren't Production
If you run a homelab Kubernetes cluster, you know it's a different beast from production. Production clusters have redundant power, dedicated hardware, and ops teams. Homelabs have Intel NUCs under a desk, electricity bills that make you think twice about running 24/7, and nodes that get rebooted — or powered off entirely — far more often than any production server should.
My cluster is three nodes: two Intel NUCs and a Proxmox VM. I shut down parts of it when not in use, it gets restarted for OS updates, and occasionally the power just goes out. The result: Kubernetes nodes go down ungracefully and come back up cold. Sometimes after such restarts not all pods come up correctly — stuck Terminating pods, Longhorn volumes refusing to attach, apps crash-looping because they started before their database was ready.
Troubleshooting these issues means running the same kubectl commands each time: checking node status, scanning pods across namespaces, verifying storage, reading logs. It's repetitive and exactly the kind of thing an AI agent should handle.
So for the Hackathon for MCP and AI Agents (MCP_HACK//26), I built PodLikar.
What Is PodLikar?
PodLikar (from Ukrainian "Лікар" — meaning "Doctor") is a read-only Kubernetes diagnostic agent built with kagent, specialised for homelab and k3s environments. It uses the Model Context Protocol (MCP) to interact with cluster tools and an LLM to reason about what it finds.
It operates in two modes:
- Health Check: A full post-reboot scan — nodes, pods, storage, dependency chains — producing a structured report with severity-classified issues.
- Targeted Diagnosis: Point it at a specific broken pod, and it gathers evidence (describe output, logs, events) to identify the root cause and suggest a fix.
The key differentiator from kagent's built-in k8s-agent is that PodLikar understands homelab-specific failure patterns — the ones that production tooling doesn't account for.
GitHub Repository: github.com/igorbnd/podlikar
Prerequisites
Before diving in, you'll need:
- A Kubernetes or k3s cluster
- kagent installed (v0.7+ tested)
- An account and API key from a supported LLM provider. kagent supports OpenAI, Anthropic, Google Gemini, Ollama, and others. I used and recommend Anthropic (Claude Haiku) — it strikes a good balance between cost and quality for this kind of task. You can get an API key at console.anthropic.com.
Architecture: No Custom Code Required
One thing that surprised me about kagent is how much you can accomplish with zero application code. PodLikar is entirely declarative — a single YAML file containing an Agent custom resource definition, a reference to a ModelConfig, a system prompt that encodes the diagnostic logic, and a list of five MCP tools from kagent's built-in tool server.
PodLikar architecture: the agent runs as a CRD, communicates with the LLM via API, and queries the cluster through MCP tools.
# podlikar-agent.yaml
apiVersion: kagent.dev/v1alpha2
kind: Agent
metadata:
name: podlikar
namespace: kagent
spec:
type: Declarative
declarative:
modelConfig: default-model-config
systemMessage: |
You are PodLikar, a specialist AI agent for diagnosing
pod health issues in homelab and k3s clusters...
tools:
- type: McpServer
mcpServer:
name: kagent-tool-server
kind: RemoteMCPServer
toolNames:
- k8s_get_resources
- k8s_get_pod_logs
- k8s_describe_resource
- k8s_get_events
- k8s_get_resource_yaml
No Python, no Go, no Docker image to build. The agent runs as a Kubernetes deployment managed by kagent's controller, communicates with the LLM via the Anthropic API, and accesses cluster data through MCP tools exposed by kagent's tool server.
Why Only 5 Tools Out of 18?
kagent's tool server exposes 18 Kubernetes tools, including write operations like ApplyManifest, DeleteResource, and PatchResource. I deliberately selected only five read-only tools:
- k8s_get_resources — list nodes, pods, PVCs, PVs
- k8s_describe_resource — detailed pod info including events
- k8s_get_pod_logs — container logs (current and previous)
- k8s_get_events — namespace events
- k8s_get_resource_yaml — full spec inspection
A diagnostic tool should never modify your cluster. Fewer tools also mean a smaller context window for the LLM, faster responses, and lower API costs.
My Cluster: The Test Environment
PodLikar runs against my 3-node K3s homelab cluster — a mix of Intel NUCs and a Proxmox VM, with Longhorn for distributed storage and a full monitoring stack. It's a proper environment running real workloads, which makes it a good testbed for diagnostic tooling.
The Prompt Engineering Journey: From 46K to 12K Tokens
This is where the real work happened, and honestly where I learned the most. I went through three major prompt iterations, each driven by testing against real failure scenarios.
The final v3 diagnostic flow: health checks use exactly 4 tool calls, targeted diagnosis uses 3.
Setting Up Test Scenarios
I created six deliberately broken pods in a podlikar-test namespace, each simulating a common homelab failure:
| Scenario | Pod Name | Failure Type |
|---|---|---|
| OOMKilled | oom-victim | Memory limit 50Mi, workload requests 128MB |
| Bad Entrypoint | bad-entrypoint | Non-existent command /bin/nonexistent-command |
| Image Pull | bad-image | Fake registry registry.example.com/totally-fake-image:v99.99 |
| Probe Failure | probe-fail | Liveness probe on port 9999, nginx listens on 80 |
| Missing Config | missing-config | References non-existent ConfigMap |
| Dependency | app-waiting-for-db | Tries to connect to PostgreSQL that doesn't exist |
These cover the core failure patterns a homelab operator would encounter. The test manifests are included in the repo.
v1: The Kitchen Sink (~46,000 tokens)
My first prompt was comprehensive — too comprehensive. It included detailed failure pattern descriptions, verbose output templates, and no constraints on tool usage.
The agent correctly identified all six broken pods. But it consumed 46,000 tokens per health check. The culprit? k8s_get_events returns the entire namespace's event history as raw JSON. For six broken pods constantly restarting, that's 20–25K tokens of event data alone.
The output was also excessively verbose — markdown tables, code blocks, multi-paragraph explanations, and a "Prevention" section nobody asked for. It even noticed the test pods had labels like scenario=oomkilled and helpfully commented that these were "deliberate test pods" — clever, but wrong behaviour for a diagnostic tool.
Key learning: An LLM will use every tool you give it and produce as much output as you let it. Constraints must be explicit.
v2: Output Formatting Rules (~30,000 tokens)
For v2, I added an "Output Rules" section with explicit constraints: no tables, no code blocks unless showing a fix, maximum 3–4 evidence bullets per pod, no "Prevention" section unless asked, and don't comment on whether pods are test pods.
I also added severity definitions (CRITICAL for crash-looping/OOMKilled, WARNING for not-ready/config errors) because v1 was classifying OOMKilled pods as mere warnings.
Better formatting, correct severity classification. But still 30K tokens because the agent was still calling k8s_get_events for every diagnosis.
Key learning: Output rules help with response quality, but the real cost is in tool calls, not output length.
v3: Tool Call Strategy (~12,000 tokens)
The breakthrough came from rethinking the tool call strategy. I added a "Token Budget Rules" section prescribing exactly which tools to call and in what order:
For health checks — exactly 4 calls:
k8s_get_resourcesfor nodesk8s_get_resourcesfor pods (all namespaces, wide output)k8s_get_resourcesfor PVCs (all namespaces)k8s_get_resourcesfor longhorn-system pods
No events. No logs. No describe. For a health check, the pod list alone tells you what's broken.
For targeted diagnosis — exactly 3 calls:
k8s_get_resourcesfor pod listk8s_describe_resourcefor the target pod (includes events!)k8s_get_pod_logswithtail_lines=20
The critical insight: k8s_describe_resource already includes recent events in its output. Calling k8s_get_events separately is almost always redundant and costs 10–20x more tokens.
I also compressed the system prompt from ~4,500 words to ~1,800 words.
Result: 12,000 tokens total. Same diagnostic accuracy:
Nodes: 3/3 Ready ✓
Storage: 8/8 PVCs Bound ✓
Pods: 86 healthy, 6 unhealthy
Issues:
CRITICAL:
- podlikar-test/oom-victim — OOMKilled, exceeded memory limit
- podlikar-test/app-waiting-for-db — cannot connect to database
WARNING:
- podlikar-test/probe-fail — liveness probe failing
- podlikar-test/bad-entrypoint — command not found
- podlikar-test/bad-image — image cannot be pulled
- podlikar-test/missing-config — ConfigMap not found
Recommended actions:
1. Fix oom-victim memory limits
2. Check database availability for app-waiting-for-db
Token Usage Comparison
| Version | Total Tokens | Tool Calls | What Changed |
|---|---|---|---|
| v1 | ~46,000 | 6–8 | No constraints, full events JSON |
| v2 | ~30,000 | 5–7 | Output rules added, still fetching events |
| v3 | ~12,000 | 3–4 | Prescribed tool sequence, dropped k8s_get_events |
The 75% reduction came almost entirely from smarter tool call strategy, not from prompt length or output formatting.
The Local LLM Experiment
A homelab agent should ideally run locally — no cloud API costs, no data leaving your network. I set up an Ollama instance on a separate VM (CPU-only, no GPU) and tested three models.
kagent makes switching trivial. You define a separate ModelConfig pointing to Ollama, and change one line in the agent:
# Cloud: Claude Haiku via Anthropic API
modelConfig: default-model-config
# Local: llama3.2 via Ollama
modelConfig: ollama-model-config
I created a lightweight variant (podlikar-local) with a ~100-word prompt, 3 tools, and a hard limit of 3 tool calls. Every model I tried — llama3.2 (3B), phi3 (3.8B), qwen2:1.5b-instruct — timed out. CPU-only inference can't process MCP tool responses fast enough for interactive use.
However, feeding pod data directly to Ollama via curl — bypassing MCP — showed the 3B model correctly identifying all six root causes in about 3 minutes:
curl -s http://192.168.8.164:11434/api/generate -d '{
"model": "llama3.2:latest",
"prompt": "Diagnose these pods: [pod list here]",
"stream": false
}' | jq -r '.response'
1. app-waiting-for-db: Insufficient or failing database connectivity.
2. bad-entrypoint: Invalid entrypoint script or command.
3. bad-image: Image pull failure due to invalid or missing image repository.
4. missing-config: Missing or incomplete configuration file.
5. oom-victim: Out of Memory (OOM) issue.
6. probe-fail: Probe (health check) failure.
The reasoning quality is there. The speed is not — at least not on CPU. With GPU acceleration or as models improve, local interactive MCP diagnostics will become practical.
Lessons Learned
- Token cost is dominated by tool responses, not prompts or output. A single
k8s_get_eventscall can consume 20K tokens. Understanding which tools are expensive is the biggest optimisation lever. k8s_describe_resourceis your best friend. It includes pod conditions, resource limits, probe configs, last state, AND recent events — all in one call.- Prescribe tool call sequences in your prompt. Don't just list available tools and hope the LLM picks the right ones. Tell it exactly which tools to call, in what order, with what parameters.
- Output formatting requires strong directive language. "Be concise" doesn't work. "NEVER use markdown tables. Maximum 3 evidence bullets. Skip preamble." does.
- Local LLMs need GPU for interactive MCP. CPU-only inference is too slow for real-time tool orchestration, but the reasoning quality is there.
- Homelab is an underserved niche. Most Kubernetes tooling assumes production clusters with ample resources and proper shutdown procedures. Homelabs have neither.
Try It Yourself
- Install kagent: kagent.dev/docs
- Get an Anthropic API key: console.anthropic.com
- Clone the repo:
git clone https://github.com/igorbnd/podlikar - Apply the agent:
kubectl apply -f podlikar-agent.yaml - Run a health check:
kagent invoke -t "health check" --agent podlikar
Test scenarios are included — deploy them and see PodLikar diagnose real failures on your own cluster.
Built for the Hackathon for MCP and AI Agents (MCP_HACK//26). PodLikar uses kagent for agent orchestration and the Model Context Protocol for Kubernetes tool access.