The problem this solves
Support teams answer the same questions every day: delivery times, return rules, account changes, fee explanations, opening hours. Customers want immediate answers and the ability to resolve simple issues themselves; agents want time for the conversations that actually need a person. Older rule-based chatbots often frustrated both, because they handled only the exact phrasings they were scripted for.
Large language models can understand varied questions and respond naturally, but used carelessly they invent policies, leak information or take actions they should not. This blueprint shows how Twara Technologies would design an assistant that is useful because it is constrained: it answers from approved content, acts only through tightly controlled tools and hands over to people at the right moments. It is a reference design, not a description of a deployed system.
Architecture
[Web / app / messaging channels]
|
[Conversation service] --- session, identity, language
|
[Input checks: abuse, PII, out-of-scope]
|
[Orchestrator] ----> [Retrieval: hybrid search over approved knowledge]
| \ ^
| \---> [Tool gateway] --> order / booking / account APIs (scoped, authenticated)
v
[LLM: generate answer with citations]
|
[Output checks: grounding, policy, format]
/ \
answer to user [Handover to human agent in helpdesk]
|
[Logs, feedback, evaluation store, analytics]
- Retrieval-augmented generation (RAG). The model is instructed to answer only from passages retrieved from the approved knowledge base, and to say so when it cannot find an answer.
- Tool gateway. Actions run through a separate service that authenticates the customer, validates parameters and enforces limits. The model can request an action; it cannot perform one directly.
- Handover. Triggers include explicit requests for a person, repeated failure to answer, negative sentiment, sensitive topics and any action above a defined risk level.
Key design decisions
Model choice. Hosted models from providers such as OpenAI, Anthropic, Google or via cloud platforms (Azure OpenAI, Amazon Bedrock, Vertex AI) offer high quality with no infrastructure to run. Open-weight models such as Llama or Mistral can be self-hosted for tighter data control but need GPU capacity and specialist operations. The blueprint keeps the model behind an abstraction so it can be swapped after evaluation, and tests candidate models against the client’s own questions, languages and tone.
Retrieval design. Pure vector search handles paraphrased questions well but can miss exact product codes or policy names; keyword search does the opposite. Hybrid retrieval (for example PostgreSQL with pgvector alongside full-text search, OpenSearch, or Azure AI Search) combines both. Passage size, metadata filters and re-ranking are tuned against a test set rather than by intuition.
Build or extend a platform. Helpdesk vendors offer AI features that work quickly within their ecosystem. A custom orchestration layer takes more effort but gives control over retrieval, guardrails, model choice and multi-channel behaviour. The blueprint suits organisations whose content, integrations or compliance needs exceed what an off-the-shelf add-on supports.
Scope of actions. Every tool the assistant can call widens what can go wrong. A typical implementation starts with read-only tools (order status, account balance summary), then adds low-risk changes with confirmation steps, and keeps refunds, cancellations of value or account-security changes with human agents.
Languages. For Indian audiences, the assistant may need to handle English, Hindi and other languages, including mixed-language messages. Model and retrieval choices are evaluated on these explicitly.
Security and compliance
The OWASP Top 10 for LLM Applications 2026 lists the main risks for systems of this kind, including prompt injection (LLM01), sensitive information disclosure (LLM02), excessive agency (LLM03), misinformation (LLM07) and hidden context exposure (LLM08). The blueprint answers each in design rather than relying on the model’s good behaviour:
- the system prompt contains no secrets, and authorisation lives in the tool gateway, not the prompt;
- tools act only on the authenticated customer’s own records, with server-side validation and rate limits;
- answers are grounded in retrieved content with citations, and output checks block unsupported policy claims;
- personal data is masked in logs where possible, and conversation retention is defined and enforced.
Conversations frequently contain personal data, so India’s Digital Personal Data Protection Act, 2023 applies. Section 5 requires a notice when consent is sought, section 8(5) requires reasonable security safeguards and section 9 sets out conditions for processing children’s data. The design includes a clear notice at the start of each conversation and configurable retention. For AI governance, it follows the Govern, Map, Measure and Manage core functions of NIST’s AI Risk Management Framework. Sector regulators may impose further requirements, which client counsel should confirm.
Phased rollout
- Discovery. Analyse support tickets and chat logs to find the most common, answerable topics; audit and clean the knowledge base.
- Internal assistant. Deploy first to support agents as an answer-drafting tool, building the evaluation set from their feedback.
- Limited public launch. Release to a subset of customers or one channel, answering questions only, with easy handover.
- Controlled actions. Add authenticated, low-risk tools with confirmations, monitored closely.
- Scale and improve. Extend channels and languages, close knowledge gaps highlighted by analytics, and re-evaluate models as options change.
Risks and how the design handles them
| Risk | How the design responds |
|---|---|
| Invented or incorrect answers | Retrieval from approved content only, citations, grounding checks and a regression test set |
| Prompt injection or jailbreaks | Input screening, no secrets in prompts, authorisation enforced outside the model |
| Assistant takes harmful actions | Narrow tools, server-side validation, confirmations and human-only high-risk actions |
| Outdated knowledge | Content ownership, automatic re-indexing on change and expiry dates on time-sensitive content |
| Customers trapped in a loop | Clear handover triggers and visible “talk to a person” option at all times |
| Model provider changes or price shifts | Model abstraction layer and portable evaluation suite |