1. End-to-End Incident Pipeline
RRT SupportOps is designed specifically to solve the bottleneck of human SRE log analysis during high-severity production outages. The system executes across five decoupled, security-hardened phases:
Local Host Ingestion & Buffer Extraction
The lightweight bash/binary agent hooks into systemd journalctl, Docker daemon events, and Nginx access/error logs. It continuously maintains a rolling 150k-token in-memory circular buffer without touching disk.
Client-Side Deterministic PII & Credential Scrubbing
Before telemetry payloads leave the host node, all bearer tokens, JWT secrets, database connection URIs, and user emails are replaced with structured hash tokens ([REDACTED_API_KEY]). Zero secrets reach the model context.
Prompt-Cached Context Dispatch to Claude 3.5 Sonnet
Ingested telemetry streams (up to 200k tokens) are dispatched to the Anthropic API. Cluster topology manifests, OS sysctl schemas, and standard operating procedures (SOPs) are marked with cache breakpoints, hitting Anthropic volatile memory.
Sandboxed Deterministic Tool Execution
When the model requests kernel or metric verifications (e.g. inspect_sysctl, query_prometheus_socket_count), the agent executes the read-only inspection strictly inside a sandboxed Linux cgroup.
Human-in-the-Loop Remediation Diff Synthesis
Rather than blindly executing modifications, the engine outputs a rollback-safe, syntax-validated unified diff file. The on-call engineer reviews and confirms before any change is applied.
2. Anthropic Prompt Caching Optimization
Processing massive 200,000-token multi-server incident traces is cost-prohibitive with standard LLM inference. RRT SupportOps leverages Anthropic's native Prompt Caching primitives to structure context hierarchically:
{
"model": "claude-3-5-sonnet-20241022",
"max_tokens": 4096,
"system": [
{
"type": "text",
"text": "<cluster_manifests>...120k tokens of topology & runbooks...</cluster_manifests>",
"cache_control": { "type": "ephemeral" } // Cache Breakpoint 1
}
],
"messages": [
{
"role": "user",
"content": "<incident_telemetry>...28k tokens live syslogs...</incident_telemetry>"
}
]
}
| Telemetry Metric | Standard LLM Ingestion | RRT SupportOps Prompt Caching | Net Efficiency |
|---|---|---|---|
| Time-to-First-Token (TTFT) | 18.4 seconds | 312 milliseconds | -98.3% Latency |
| Input Token Cost (100k tokens) | $0.30 / request | $0.03 / cache hit | -90.0% Operational Cost |
| Max Ingestion Depth | 16k tokens truncated | 200k tokens full trace | Zero Context Loss |
3. Agent Configuration Reference (`agent.yaml`)
The host agent is configured via a standard YAML manifest located at /etc/rrt/agent.yaml:
# RRT SupportOps Node Agent Specification
version: "1.2"
node:
id: "node-sg-01.rrt.internal"
region: "ap-southeast-1a"
environment: "production"
telemetry:
gateway_endpoint: "https://rrt.vn"
buffer_window_tokens: 150000
flush_interval_seconds: 5
sources:
journalctl:
enabled: true
units: ["nginx.service", "docker.service", "sshd.service"]
files:
- path: "/var/log/nginx/error.log"
format: "nginx_error"
- path: "/var/log/syslog"
format: "syslog"
security:
pii_redaction: true
strip_authorization_headers: true
enforce_read_only_tools: true
allowed_tool_syscalls: ["inspect_sysctl", "query_prometheus", "read_journalctl"]
4. Deterministic Sandboxed Tool Whitelist
To eliminate hallucination risks, Claude 3.5 Sonnet calls strictly registered, read-only tools:
| Tool Identifier | Execution Scope | Safety Policy |
|---|---|---|
inspect_sysctl_net_core() |
Reads somaxconn, tcp_max_syn_backlog, tcp_tw_reuse |
Read-Only |
query_prometheus_socket_count() |
Queries active TIME_WAIT, ESTABLISHED, and SYN_RECV metrics |
Read-Only |
check_cgroup_memory_pressure() |
Scrapes /sys/fs/cgroup/memory.pressure and PSI indicators |
Read-Only |
verify_journalctl_errors() |
Filters kernel panics, OOM triggers, and segmentation faults | Read-Only |
5. Verified Production Postmortems
Case #1041: TCP Socket Starvation under 10Gbps Ingress Surge
Context: A high-concurrency microservice cluster behind Nginx began dropping connections with intermittent HTTP 502 Bad Gateway responses during a traffic peak. Host CPU was under 18%.
Claude 3.5 Sonnet Diagnosis: Ingesting 148,000 tokens of Nginx error buffers and kernel syslogs, the model correlated TCP: request_sock_TCP: Possible SYN flooding with worker_connections are not enough. It identified that net.core.somaxconn=128 (Linux default) truncated the TCP backlog, while net.ipv4.tcp_tw_reuse=0 exhausted ephemeral ports across NAT boundaries.
Remediation: Synthesized a targeted diff for /etc/sysctl.d/99-network-tuning.conf setting somaxconn=65535 and tcp_tw_reuse=1. MTTR dropped from 48 minutes to 2 minutes 15 seconds.
Case #1042: Cascade Docker Cgroup v2 OOM Killer Remediation
Context: A background telemetry ingestion container was killed intermittently without generating fatal application tracebacks.
Claude 3.5 Sonnet Diagnosis: Evaluated container_oom: cgroup /docker/d3b09 memory limit reached against host Prometheus RSS memory allocations. Discovered an unconstrained Python circular reference accumulating memory in worker sub-processes competing for a strict 1024MB cgroup cap.
Remediation: Synthesized a Docker Compose memory limit adjustment with soft limits (mem_reservation: 768m, mem_limit: 2048m) and proposed a targeted garbage collection runbook.
6. Early Access Pilot SLA & Requirements
Qualified enterprise DevOps and infrastructure engineering teams can enroll in the 60-day RRT SupportOps Pilot program:
| Pilot Dimension | Specification | Governance Guarantee |
|---|---|---|
| Node Quota | Up to 50 Linux servers / containers | Dedicated telemetry quota |
| Diagnosis SLA | < 15 seconds mean analysis time | High-priority rate limits |
| Data Retention | Zero disk persistence on RRT gateways | GDPR & SOC2 Type II compliant |
| Engineer Support | Direct escalation to Founder / Principal Architect | 24/7 incident triage channel |