Skip to main content

Sentinel

Production AI ops layer on AWS Bedrock — investigate-only Claude agents, Slack-gated escalation.

AI Infrastructure·Beta·Rev. 2026·AWS Bedrock · Claude Sonnet 5.5 · Slack

What is Sentinel?

Sentinel is iSimplifyMe's production AI operations layer — a fleet of investigate-only Claude agents running on AWS Bedrock that monitor iSM's own infrastructure for regressions, anomalies, and operational incidents. Three workloads are deployed in production: a Diagnostics Agent for tenant uptime forensics, a GH Triage Agent for CI failure classification, and a Pipeline Hang Detector for content-pipeline anomaly investigation. Every detection fires into Slack with structured context, a recommended action, and a human approval gate before any remediation runs.

Abstract

Sentinel is iSimplifyMe's production AI operations layer — a fleet of investigate-only Claude agents running on AWS Bedrock that monitor iSM's own infrastructure for regressions, anomalies, and operational incidents. It is internal infrastructure, not a customer-facing product, and runs to keep the rest of the platform honest. The same architecture is offered to clients as a productized Sentinel-pattern monitoring retainer.

Problem

Production AI infrastructure has more silent failure modes than monitorable ones. A Bedrock model that responds with semantically wrong answers passes a 200 OK health check. A retrieval pipeline that surfaces stale data clears every uptime probe.

Manual log review does not scale across a 31-tenant fleet. Status pages tell you what is on; they do not tell you what is wrong.

Approach

The agent topology

Each Sentinel agent is an investigate-only Bedrock-hosted workload with a discrete surveillance scope and a defined cadence. The agent reads from a constrained set of operational signals — logs, recent error events, model-call traces — runs a Claude Sonnet 5.5 inference pass via the us. US-bounded inference profile to classify the situation, and decides whether the finding warrants escalation. No agent writes to client systems; the architecture is investigate-and-notify only.

Slack as the approval gate

When an agent identifies something worth escalating, it posts a Block Kit card to the appropriate channel with the diagnosis, the recommended remediation, and a small set of action buttons. A human reviewer clicks one. Only then does any remediation fire.

Sentinel's design rule: automated detection is fine; automated remediation requires a human in the loop. An agent can find a problem at any hour, but only a person decides what changes in response.

Workload #1: Diagnostics Agent

The Diagnostics Agent investigates client tenant sites that have failed three consecutive uptime checks — running curl, dig, and Cloudflare 5xx-breakdown probes via custom Bedrock tools — then files a markdown bug-report ticket with timeline, root cause, evidence, and recommended fix. Verified cost: $0.0822 per incident on synthetic test cases after the May 2026 Bedrock migration, measured on Claude Sonnet 4.6; the workload has run on Claude Sonnet 5.5 since October 2026.

The full path is deployed in production with per-tenant activation canary-gated; the workload has its own dossier at /labs/diagnostics-agent.

Workload #2: GH Triage Agent

The GH Triage Agent polls workflow runs across the 82 iSimplifyMe org repositories every fifteen minutes, detects failures, and runs an inference pass classifying root cause across eight categories: test_flake, regression, infrastructure, auth, dependency, lint_or_typecheck, build_config, and unknown. Output is a structured ticket with markdown body covering Failure Summary, Classification, Failed Jobs, Recent Commits, and investigator Notes. Verified cost: $0.065 per run, measured on Claude Sonnet 4.6 at migration.

Idempotent — once a failed run is investigated, a 24-hour DDB lock prevents re-investigation, so flapping CI does not produce duplicate tickets.

Workload #3: Pipeline Hang Detector

The Pipeline Hang Detector watches the iSM multi-site content pipeline for anomaly states — stuck topic-proposal runs, malformed MDX rejections, frontmatter envelope drift, write-post Lambda failures — and investigates each one, classifying the cause and identifying the affected tenants. Output uses the same structured ticket format as the other agents and routes to a content-pipeline-specific Slack channel.

The three workloads share infrastructure: one generic SQS-triggered runner Lambda dispatches the right agent based on a SENTINEL_AGENT_SLUG kickoff message, an atomic conditional-write lock at INCIDENT#OPEN race-protects parallel detection paths, and the same file_ticket and notify_slack tools serve all three. Adding a new Sentinel workload is a registry entry plus a detector handler; everything else is shared.

The same substrate as client builds

Sentinel runs on the AWS Bedrock substrate (BedrockRuntimeClient + ConverseStreamCommand + DynamoDB ticket store + EventBridge cron + SQS queue + IAM-scoped Bedrock perms) that iSimplifyMe deploys for client validator-architecture engagements. iSM operates Sentinel as the production proof point of the architecture it proposes for regulated-industry clients.

Status

  • Sentinel runs in production on AWS Bedrock as iSM's internal AI operations infrastructure. Three workloads are deployed: Diagnostics Agent (per-tenant activation canary-gated), GH Triage Agent, and Pipeline Hang Detector.
  • Total platform cost: under $50/month across all three workloads at current activity volume.
  • Every agent is investigate-only by design. None writes to client systems, and no remediation fires without human approval through the Slack gate.

For client engagements

Sentinel-pattern monitoring is a monthly retainer that runs the same architecture on the client's own AWS Bedrock account. Its investigate-only agents watch validator-gate hits and misses, drift, and incidents.

Findings go to a Slack channel the client names, and a person approves any remediation before it runs. iSM reviews each deployment every quarter against its reference architecture.

The retainer follows the Validator Architecture build and is sized to scope. It is open only to mid-market and enterprise companies in regulated industries, and engagements start with the Validator Gap Audit.

Frequently asked

Apex Architecture

Every site we build runs on Apex — sub-500ms, AI-native, zero maintenance.

Explore Apex Architecture

Stay Ahead of the Curve

AI strategies, case studies & industry insights — delivered monthly.

⌘ K