Skip to main content

AI Concierge

One production AI route serving every fleet widget: per-tenant isolation, budgets, guardrails.

AI Infrastructure·Shipped·Rev. 2026·Bedrock · S3 · DynamoDB

What is the iSimplifyMe AI Concierge?

The iSimplifyMe AI Concierge is a production multi-tenant AI system: one Bedrock-backed API route serving every concierge widget across the iSM network from per-tenant configuration. Each tenant carries its own persona, knowledge base, monthly usage budget, and escalation rules. As of September 2026, the fleet runs twenty-eight live tenant configurations, and the medical flag with stored-record sanitization is set on eight production tenants. Every concierge lead carries the fleet's standard first-touch attribution.

Abstract

The AI Concierge is one production AI system operated across the iSimplifyMe network: a single Bedrock-backed route that serves every tenant widget in the fleet — including regulated medical practices — from per-tenant configuration. Each tenant carries its own persona, knowledge base, usage budget, and escalation rules on shared, metered infrastructure. This dossier documents the isolation model, the guardrails, and the operating numbers.

Problem

Embedding a chat widget is the easy part. Operating a conversational AI surface across a fleet of properties — where one tenant is a medical practice with strict retention obligations and the next is a retail brand — is an infrastructure problem: isolation, spend control, escalation, and attribution all have to hold per tenant, on shared plumbing.

Off-the-shelf tools solve the widget and stop there. The operating layer is what decides whether the system can sit anywhere near a regulated client.

Approach

One route, resolved per tenant

Every widget in the fleet calls the same central API route. That route resolves the tenant from the request's domain header, loads the tenant's knowledge.json and persona.json from a tenant-scoped S3 prefix, and assembles the system prompt server-side, per request, before any model call.

Once the tenant resolves, its files are schema-validated on load and cached in memory for five minutes. Sessions live in DynamoDB with a thirty-minute TTL, responses stream to the widget over SSE, and inference runs on Claude via AWS Bedrock.

Input discipline

A regex input classifier screens every visitor message before it reaches the model: sixteen documented injection and jailbreak patterns, a 500-character ceiling, and a control-character screen. An injection_resistance block is prepended to every tenant system prompt.

Usage governance

The public widget is capped at 1,000 messages per tenant per month, enforced with atomic counters and Slack alerts at 50, 75, and 90 percent of budget. A per-visitor throttle allows twenty messages a day, keyed to an HMAC-hashed IP that exists only to correlate one visitor's activity within a day. Caps fail open on infrastructure errors, so a database fault cannot take a tenant's widget down.

Regulated tenants

Tenant profiles carry a medical flag; eight production tenants run with it set. For flagged tenants, the multi-tenant lead store drops the free-text fields that can carry health narratives — the conversation transcript and arbitrary custom fields — from the stored record, while structured contact identity is retained so the lead stays actionable. The practice's own notification email still receives the full submission.

Escalation is configured per tenant. Each persona can define emergency signature phrases the model emits when a conversation reads as urgent; the route scans the reply stream and, on first match, sends a structured emergency event that pins the tenant's human-contact action in the widget before the reply finishes rendering.

Attribution, end to end

Concierge leads carry the same first-touch attribution as every other lead in the fleet. A pixel writes a thirty-day first-touch cookie with UTM, click-ID, AI-referrer, and referrer precedence; a CloudFront edge function stamps a one-hour AI-referrer cookie for eight AI answer engines.

Both cookies are first-party on the client's own domain, and the chat call reaches Apex through the site's server-side proxy, so the widget reads them in the browser and sends them in the chat body. Apex resolves them into the same attribution fields the form pipeline writes.

Status

Everything below is verified against production code and live store aggregates, as of 2026-09-26:
  • The per-tenant store holds twenty-eight live tenant configurations, and eight production tenants carry the medical flag.
  • Excluding records flagged as spam, the shared lead register has captured 470 leads across twenty-two tenants since April 2026. Of those, 117 arrived through concierge surfaces — including leads produced by the urgent and emergency capture paths.
The concierge is operating proof of the same pattern iSimplifyMe builds for clients: shared AI infrastructure with tenant-scoped trust boundaries and budgets that alarm before they surprise. Sentinel watches the fleet the same way; client builds of the pattern run through AI agent architecture.

Links

Frequently asked

Apex Architecture

Every site we build runs on Apex — sub-500ms, AI-native, zero maintenance.

Explore Apex Architecture

Stay Ahead of the Curve

AI strategies, case studies & industry insights — delivered monthly.

⌘ K