Skip to main content

66 posts tagged with "Architecture"

Cross-cloud AI architecture decisions and patterns

View All Tags

Amazon Nova and Bedrock for Engineers: When AWS-Native Actually Beats Best-of-Breed

· 11 min read
AI Playbook author

Amazon Nova is easy to underrate if you only compare it on public leaderboards against GPT-5.6 or Claude 5 — it isn't trying to win that fight. Nova's actual pitch is that if you're already an AWS shop, the model, the customisation tooling, the agent framework and the deployment surface are the same product, with the same IAM, the same billing, and the same on-call rotation. That's a genuinely different value proposition, and it's the one you should evaluate honestly rather than dismissing Nova as "the AWS also-ran."

Claude 5 for Engineers: Fable, Opus and Sonnet — Reading the Tiers, the Pricing Clock and MCP

· 11 min read
AI Playbook author

Anthropic's Claude 5 launch is really three separate launches with one shared story: capability is moving down the price ladder faster than most teams' architecture assumes. Opus 5 landed on 24 July 2026 at the exact same price as the outgoing Opus 4.8 — $5 input / $25 output per million tokens — while offering Fable-adjacent performance on several benchmarks. If your Claude cost model still assumes "the good model is expensive," it's already out of date.

Command A+ for Engineers: The Sovereign AI Argument, Made With an Apache Licence

· 11 min read
AI Playbook author

"Sovereign AI" gets used loosely enough in vendor marketing that it's worth being precise about what Cohere's Command A+ actually offers versus what the phrase implies. Command A+ (218B and 25B, Apache 2.0) is built specifically for enterprises and governments that need to run a capable model entirely within their own infrastructure boundary — not as a philosophical statement, but as a concrete set of deployment, data-residency, and control guarantees that a hosted API fundamentally cannot provide, regardless of that API vendor's compliance certifications.

DeepSeek for Engineers: MLA, MoE, and GRPO — the Three Ideas Behind the Cheapest Frontier-Class Training Run

· 11 min read
AI Playbook author

DeepSeek's real contribution to the 2026 model landscape isn't any single model — it's three specific technical ideas that the rest of the industry has since absorbed to varying degrees: Multi-head Latent Attention for cheaper KV-cache memory, Mixture-of-Experts done at extreme sparsity, and Group Relative Policy Optimisation as a reinforcement-learning method that made reasoning-focused training dramatically cheaper. If you understand those three ideas, you understand why DeepSeek mattered, and you'll recognise the same ideas (in variant form) inside half the other models in this series.

From Chatbots to Autonomous Digital Workers: How Frontier AI Models Are Becoming Smarter Over Time

· 33 min read
AI Playbook author

Artificial intelligence models are improving at a remarkable pace. A model released only six months ago can quickly appear less capable, less efficient and less reliable than a newer generation. It is tempting to assume that every new model is smarter because it contains more parameters — but modern frontier systems such as Claude Fable 5, Claude Opus 5, GPT-5.6, Gemini 3.6 Flash and Kimi K3 are improving through a much broader stack of advances.