Case study · trynowreply.com · Live
nowreply
An AI front desk for appointment-based businesses. Connect your documents, website and calendar; the assistant answers every inquiry in seconds from your own prices and policies, on your website and in your inbox, and books the appointment. Underneath it is starmo's multi-tenant retrieval platform, built for grounded, cited answers.
The product demos itself. The chat on the landing page is the production assistant, not a mockup. What a visitor watches it do there, answer from documents, cite them, decline to invent, is what it does on a customer's site and in their inbox.
The problem
Every inquiry is a question about prices, policies or a free slot
A med-spa, a salon, a clinic: the inquiries that arrive on the website and in the inbox are mostly the same few questions. What does this cost, what does the treatment involve, what is the cancellation policy, is there anything on Thursday. Each one is small. Together they are a front desk's day, and at a small business the front desk is often the owner, between clients.
The answers exist. They are in the price list, the policy page and the calendar. What is missing is someone available to give them in the minute the customer is asking, and then to take the booking. A generic chatbot fails here in a specific way: it improvises. For a customer-facing agent that quotes prices and books appointments, an invented policy is worse than no answer.
So the product had to do three things: answer only from the business's own material, say so when that material does not cover the question, and turn the answered question into a booked appointment inside the same conversation.
The approach
Four decisions
Grounded, or it says so
Every answer comes from retrieved documents, with inline citations. When the documents do not cover a question, the assistant says exactly that and offers what they do cover, rather than inventing a policy.
Knowledge that maintains itself
Files, the business's website and Google Drive folders are ingested once and kept current by scheduled re-crawls and sync. Live systems join through OpenAPI, so an endpoint becomes a tool the agent can call mid-answer. Nobody maintains an FAQ.
One agent, every channel
The same agent answers on the website widget and in the email inbox, and books through Google Calendar inside the conversation. Persona and behaviour are per-business configuration, not code.
Every conversation is a trace
Duration, tokens and cost are logged per query, with filters by confidence, channel, outcome and feedback, and a trace inspector for stepping through what the agent did. Tenant isolation is asserted by dedicated test suites, not assumed.
What shipped
The public page, then the workspace
The screens below are the live product: the landing page first, then the workspace of a demo med-spa account.

The demo is the product. The landing page embeds the production assistant. A visitor asks it something and watches it answer from the demo business's documents, which makes the case better than a description of the feature would.

Three kinds of source. Uploaded files are chunked and embedded; websites are crawled and re-crawled on a schedule; live-data APIs are imported from an OpenAPI spec and become tools the agent can call during a conversation, shared by every agent in the workspace. Google Drive folders sync on their own.

Answers that show their work. A pricing question in the workspace chat: the grounded answer streams back with inline citations, the retrieval trace is one click away (sources searched, one found, 2.0 seconds), and the source document renders beside the reply. A business can verify any answer its front desk gives.

“I don't know”, on purpose. Asked whether Botox can be combined with a HydraFacial, the assistant answers the documented part with a citation and says plainly that the documents do not state the rest. For an agent that quotes prices and books appointments, refusing to improvise is the product.

Configured, not coded. Display name, greeting, behavioural instructions, generated starter questions, theme and accent colour, with a live preview of the widget as it will render on the business's own site, updating as they type.

Website, inbox, calendar. Add website chat, or connect an email inbox so inquiries by mail get the same grounded answers. Connect Google Calendar once and every agent can check real availability and book the appointment in conversation: the moment an answered question becomes revenue.

An engineering system, reviewed like one. Every exchange is logged with its duration, token count and dollar cost, filterable by confidence, channel, outcome, status and user feedback. The conversations on screen cost between $0.0007 and $0.0028 each.
The platform
A multi-tenant retrieval spine
The retrieval spine has the same shape as the one in DocsAI, built from scratch with no framework, and hardened here for many tenants on one deployment. Every organization's knowledge, agents and conversations are its own; the stages are shared.
Blue is the answer path after retrieval: the hybrid search, the streamed reply, and the tools it may call on the way. The dashed line is the trace every conversation leaves. The two zones in the middle are what a tenant owns; the spine is what they share.
Tenant isolation, proven
Organizations, roles and invitations, with dedicated end-to-end suites asserting that one organization can never read another's data. Isolation is tested, not assumed.
Pluggable verticals
Domain tool packs register typed tool specifications into a shared registry and are activated per deployment by environment variable, so the core stays industry-agnostic. Persona lives in per-organization system prompts and glossaries: tenant data, not code.
Tools from OpenAPI
An importer turns an OpenAPI spec into callable agent tools with citation handling, so a live system becomes knowledge without custom integration code. Tools are workspace-wide and shared by every agent.
Knowledge that maintains itself
Scheduled website re-crawls and Google Drive sync keep each corpus current; documents version through an explicit ingestion pipeline.
Voice-ready
Live voice sessions with transcription logging are built in as an optional mode: the same grounded core, spoken.
Shipping it
The layer most prototypes never reach
Tests. 755 backend tests across 66 files run in CI against a real pgvector Postgres, plus Playwright end-to-end suites for authentication, role-based access, invitations and cross-tenant boundaries.
Migrations. The schema has evolved through 46 migrations without a reset, which is the unglamorous proof of a system in operation rather than one rebuilt for each demo.
Billing and onboarding. Paddle subscriptions, a guided five-step setup (create the agent, teach it, try it in chat, install the script tag, watch the first visitor conversation arrive), usage metering and audit logging: everything between “it works” and “it's a business”.
Deploy. FastAPI and the built React frontend ship as one Docker image to Railway. Uploads go straight to MinIO or S3 through presigned URLs, so files never proxy through the API.
Stack
What it is made of
| Layer | Choice |
|---|---|
| Platform | Multi-tenant retrieval platform: FastAPI, Postgres with pgvector, organizations, roles and invitations |
| Retrieval | Router, rewriter, expander, hybrid pgvector and full-text search, reranker, retrieval evaluator; streamed over Server-Sent Events with inline citations |
| Knowledge | Uploaded files, website crawl with scheduled re-sync, Google Drive sync, tools imported from OpenAPI specs |
| Channels | Website widget installed with one script tag and per-agent signed keys, email inbox, Google Calendar booking, optional live voice |
| Frontend | React 19, Vite, Tailwind |
| Billing | Paddle subscriptions, usage metering, audit logging |
| Tests | 755 backend across 66 files, in CI against real pgvector; Playwright end-to-end including tenant boundaries |
| Scale | About 72,000 lines, 46 migrations |
| Deploy | Railway, one Docker image; MinIO or S3 presigned uploads |