Skip to main content

66 posts tagged with "Architecture"

Cross-cloud AI architecture decisions and patterns

View All Tags

LLM Cost Tracking and FinOps Across Azure, AWS, and Google Cloud

· 33 min read
AI Playbook author

Generative AI cost management is more complicated than ordinary cloud cost management. A conventional application might be charged according to CPU hours, memory, storage, and network usage. An LLM application can accumulate costs through input tokens, output tokens, cached tokens, reasoning tokens, embeddings, retrieval, reranking, guardrails, model evaluation, agent tool calls, retries, observability, and the infrastructure supporting the application.

The main FinOps challenge is therefore not simply:

How can we use the cheapest model?

The correct question is:

What is the lowest sustainable cost at which the application can meet its quality, latency, reliability, security, and business-outcome requirements?

LLM Response Caching Strategies for RAG and Agentic Applications

· 30 min read
AI Playbook author

Caching in a conventional application usually means storing the result of a database query or API request. In an LLM application, there are many more opportunities: repeated questions, equivalent paraphrases, identical embeddings, repeated retrieval and reranking, repeated tool calls, stable system prompts, and recurring agent planning steps.

The strongest architecture therefore does not implement one “LLM cache.” It implements a hierarchy of specialised caches across the RAG and agent pipeline.

Security, Compliance and Governance for Open-Source and Closed-Source LLM Deployments

· 39 min read
AI Playbook author

Deploying a large language model is not simply a question of choosing between an open-source model and a commercial API. It is an enterprise risk decision involving:

  • What information the system will process.
  • Where that information will travel.
  • Who can access the model, prompts, outputs and logs.
  • What actions the model can perform.
  • How the organisation will detect failures or attacks.
  • Which party is accountable when something goes wrong.
  • What evidence can be presented to auditors, regulators, customers and executives.

Production-Grade Evaluation Strategies for LLMs, RAG, Agents, and Multi-Agent Systems

· 41 min read
AI Playbook author

Evaluating a conventional machine-learning model is often straightforward. A classification model can be measured using accuracy, precision, recall, F1 score, or area under the curve. A regression model can be measured using mean absolute error or root mean squared error.

Large language model applications are more difficult to evaluate because they are:

  • Non-deterministic
  • Generative rather than strictly predictive
  • Capable of producing several valid answers
  • Often composed of multiple models, tools, retrievers, databases, prompts, and agents
  • Able to modify external environments
  • Sensitive to prompt wording, model versions, context ordering, retrieved documents, and tool responses
  • Expected to satisfy qualitative requirements such as helpfulness, clarity, faithfulness, tone, safety, and policy compliance

A production-grade evaluation strategy therefore cannot rely on one benchmark or one quality score. It must evaluate the system at several levels:

  1. The underlying model
  2. Individual LLM responses
  3. Retrieval and generation components
  4. Tool calls and agent trajectories
  5. Multi-agent coordination
  6. End-to-end business outcomes
  7. Operational performance
  8. Safety, compliance, and governance
  9. Real production behaviour

Hugging Face Evaluate provides useful building blocks for metrics, comparisons, and measurements. However, Hugging Face currently presents LightEval as the more actively maintained toolkit for modern LLM benchmarking, while agentic applications require additional trace, tool, environment, and outcome evaluation capabilities.

The Integrated 8D AI Solution Engineering Framework: Banking Customer-Service Final Playbook

· 15 min read
AI Playbook author

The 8D AI Solution Engineering Framework turns an unclear AI ambition into a valuable, secure, governed and operational service. For MonGo Bank, it transforms “build a chatbot to cut cost” into a trusted hybrid customer-service capability—and maps every framework from Parts I–VIII into one controlled learning cycle.

End-to-End AI Solution Engineering Playbook: Architecture, Operating Model and Engineering Design for Banking Customer Service

· 14 min read
AI Playbook author

A funded hybrid AI programme still fails if MonGo ships a strong model inside a weak system. Architecture and operating design must cover channels, authentication, banking APIs, knowledge, retrieval, models, guardrails, evaluation, escalation, monitoring, governance, cost and ownership.

This article is Part IV of the Banking Customer-Service AI playbook. It follows Part I, Part II and Part III.

End-to-End AI Solution Engineering Playbook: Commercial Case, Benefits and Investment for Banking Customer Service

· 15 min read
AI Playbook author

Strategic fit and readiness do not fund a programme. MonGo Bank must still prove what the hybrid AI service will cost, which benefits are cash versus capacity, who owns them and when to continue, expand or stop.

This article is Part III of the Banking Customer-Service AI playbook: commercial case, benefits and investment. It follows Part I: Strategy and Discovery and Part II: Readiness, Maturity and Prioritisation.

End-to-End AI Solution Engineering Playbook: Delivery, Change, Adoption and Operations for Banking Customer Service

· 13 min read
AI Playbook author

A approved, evaluated AI system still fails if employees distrust it, managers keep old metrics, operations lack ownership or benefits never convert to value. Delivery means establishing a reliable AI-enabled service people use correctly—not merely deploying a model.

This article is Part VII of the Banking Customer-Service AI playbook. It follows Part I through Part VI.

End-to-End AI Solution Engineering Playbook: AI Engineering, Evaluation and Experimentation for Banking Customer Service

· 14 min read
AI Playbook author

An AI demo proves a model can produce an answer. AI engineering proves the complete system can produce acceptable outcomes repeatedly, safely and economically—before MonGo exposes it to customers and employees.

This article is Part V of the Banking Customer-Service AI playbook. It follows Part I through Part IV.

End-to-End AI Solution Engineering Playbook: Portfolio Scaling, Enterprise Transformation and Continuous Value

· 11 min read
AI Playbook author

One successful customer-service AI product answers “can we build something useful?” The enterprise question is whether MonGo can scale AI across products and functions without duplicated platforms, inconsistent controls, uncontrolled cost or fragmented ownership.

This article is Part VIII of the Banking Customer-Service AI playbook. It follows Part I through Part VII.