The way in opens the moment the set has loaded.

Architecture of IntelligenceDrawing set · Akshay Bajpai
The complete set35 sheets in one document. Choose “Save as PDF” in the print dialog.

G-000 The Architecture of Intelligence

Akshay Bajpai

Architect of systems · Builder of intelligence

Document
Drawing set, complete
Sheets
35
Issued
2026.10
Drawn
A. Bajpai
Live at
www.akshaybajpai.com

Drawing index

General arrangement

  1. G-000Key Plan
  2. A-101The Architect
  3. R-301Research
  4. S-201Structural Principles
  5. W-400Works
  6. E-600Essays
  7. B-500Field Notes
  8. C-700Correspondence

Works · details

  1. W-401Forward-Deployed: Multi-Tenant Fertility & Surrogacy AI2026.05
  2. W-402AIDA: Health Programme Intelligence You Can Brief From2026.04
  3. W-403Agentic Video Intelligence for 24/7 Operations2026.04
  4. W-404Neural Map Personal Site: Spatial Navigation as Homepage2026.02
  5. W-405AURIXA: Conversational AI Orchestration at Scale2026.02
  6. W-406Alzheimer's Classification: Master's Thesis Research2026.01
  7. W-407Agentic Finance & Multimodal Healthcare AI2025.11
  8. W-408Publishing AI Automation: Multi-Channel Support at Scale2025.09
  9. W-409Restaurant SaaS Forecasting at Scale2025.03
  10. W-410Insurance Document Intelligence on AWS2024.06
  11. W-411Clinical NLP from Unstructured EHRs2022.06

Essays · details

  1. E-601The Model That Refuses to Draw2026.10
  2. E-602Minimalism as Engineering2025.03
  3. E-603The Architecture of Trust2025.02
  4. E-604Why Performance Is a Feature2025.02
  5. E-605The Cost of Convenience2025.01
  6. E-606What Systems Thinking Actually Is2025.01

Field notes · details

  1. B-501Reinforcement Learning After the Reward Model2026.10
  2. B-502World Models, As Issued: The State of the Set in October 20262026.10
  3. B-503Experimental UI as an Engineering Discipline2025.03
  4. B-504Architecture Design for Healthcare AI2025.03
  5. B-505AI in Restaurant Automation2025.02
  6. B-506Systems Thinking for Founders2025.02
  7. B-507Performance Engineering on Static Hosting2025.02
  8. B-508Building SaaS with a Zero-Dependency Mindset2025.02
  9. B-509DNS-Level Ad Blocking. System Design2025.01
  10. B-510AI Infrastructure Philosophy2025.01

A-101Architectural

The Architect

Biography, trajectory, and operating principles

Architect of systems. Builder of intelligence.

Scale
1:1
Rev
C
Issued
2026.09

I work where product stakes, model behavior, and infrastructure meet. The through-line in my career is taking ambiguous domain problems (multi-tenant fertility platforms, defense-adjacent video intelligence, and clinical NLP) and making them operable.

Technically, I live in the stack you actually run in production: LLMs and SLMs, agentic orchestration, hybrid RAG, schema-grounded NL2SQL, FastAPI, React, Kafka, Docker, and AWS CDK. I care about evaluation, cost-aware routing, and the boring parts: JWT auth, audit logs, and Playwright regression. That is what separates a demo from something a C-suite can sign off on.

Chronology

2021202320252026
  1. 2021–2023

    MSc Artificial Intelligence

    Lviv Polytechnic National University

    Graduated with Distinction (9.8/10). Master's thesis on Alzheimer's disease classification benchmarking 9 machine learning models on longitudinal biomarkers.

  2. 2023–2025

    Data Scientist II

    Insurance · Document Intelligence

    Shipped document intelligence on AWS Bedrock (Claude, LayoutLMv3) achieving 95%+ extraction accuracy. Built ensemble underwriting models improving efficiency by 78%.

  3. 2025

    Founding Engineer

    Restaurant SaaS

    Validated demand forecasting across 200+ pilot sites using XGBoost and Prophet, reducing food waste by 32%. Scaled event-driven pipelines to process 2M+ events/day.

  4. 2025–2026

    Senior AI Consultant

    Finance & Healthcare

    Built stateful LangGraph finance workflows reducing manual touchpoints by 60%. Deployed multimodal clinical pipelines (QLoRA, hybrid RAG) on HIPAA-aware AWS, maintaining sub-800ms p95 latency.

  5. 2026

    AI Lead

    Defense-Adjacent · Video Intelligence

    Architected an agentic video-intelligence platform over 24/7 CCTV using LangGraph and MCP, cutting analyst intervention by ~70%. Engineered Kafka ingestion for sub-200ms latency on air-gapped infrastructure.

  6. 2026–Present

    Lead Full-Stack AI Engineer

    Forward Deployment · Multi-Tenant AI

    Owning roadmap-to-production across multiple client programs. Architecting shared LLM gateways, governed reporting agents, and Python microservices on AWS CDK. Delivering multi-tenant workflows that survive C-suite UAT.

Fig. 1Elevation along the career datum. Choose a volume to read the role; each stands on the one before.

This site is a thinking laboratory: case studies, technical writing, and essays that mirror how I negotiate scope, architecture, and delivery. Organization names are generalized on this site; the engineering is specific. For AI systems, forward-deployed engineering, or performance-critical products, get in touch.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleThe ArchitectRevCIssued2026.09A-101

A-101record

sheet
A-101
title
The Architect
subtitle
Biography, trajectory, and operating principles
discipline
A · Architectural
scale
1:1
revision
C
issued
2026.09
refs
R-301, W-400, C-700
name
Akshay Bajpai
role
Senior Full Stack AI Engineer · Forward Deployment
based
New Delhi, India

R-301Research

Research

Published work and experimental directions

Published work and experimental directions.

Scale
1:1
Rev
B
Issued
2026.09
Degree
MSc AI, 9.8/10
Publication
Springer 2024
1Benchmark plot, nine modelsR-301 · 1:1

My formal training is in intelligent systems and medical AI. I completed a Master of Science in Artificial Intelligence & Intelligent Systems at Lviv Polytechnic National University (2021–2023) with a CGPA of 9.8/10 and distinction. Before that, a B.Tech in Computer Science & Engineering from Rajiv Gandhi Prodyogiki Vishwavidyalaya, Bhopal (2017–2021), with honours.

Schedule of record5 entries
  1. 2024Publication · lead authorChapter 17: machine learning approaches to medical diagnosisIn ML for Medical Diagnosis in Data-Centric Business and Application, 3rd edition. Springer, ISBN 978-3-031-60815-5.
  2. 2021–2023DegreeMSc, Artificial Intelligence & Intelligent SystemsLviv Polytechnic National University. Awarded with distinction.9.8/10
  3. 2021–2023Master's thesisAlzheimer's classification on OASIS biomarkersNine models benchmarked with explicit preprocessing choices and recall-weighted evaluation.9 models
  4. UndergraduateResearch studyComparative study of Alzheimer's diagnosis using machine learningAuthored at IIT Gandhinagar. The methodological foundation for the thesis.
  5. 2017–2021DegreeB.Tech, Computer Science & EngineeringRajiv Gandhi Prodyogiki Vishwavidyalaya, Bhopal. Awarded with honours.

Peer-reviewed publication. I am lead author on Chapter 17, machine learning approaches to medical diagnosis, in ML for Medical Diagnosis in Data-Centric Business and Application, 3rd edition (Springer, 2024, ISBN 978-3-031-60815-5). The chapter situates diagnostic models inside data-centric business constraints: label quality, deployment accountability, and the gap between benchmark accuracy and clinical utility.

Undergraduate research. At IIT Gandhinagar I authored a comparative study on Alzheimer's disease diagnosis using machine learning, the methodological foundation for my later Master's thesis on OASIS biomarkers, where nine models were benchmarked with explicit preprocessing choices and recall-weighted evaluation.

Production work since then (insurance document intelligence, EHR variable extraction, multimodal diagnostic imaging, programme KPI engines) extends the same principle: rigor in data, honest metrics, and systems that clinicians and operators can override. Deeper build narratives live under Work; opinion and infrastructure philosophy under Blog and Essays.

Current research interests include governed agentic retrieval, schema-grounded text-to-SQL, minimal-dependency edge deployments, and evaluation pipelines that survive executive readouts, the same problems I ship against in forward-deployed engagements.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleResearchRevBIssued2026.09R-301

R-301record

sheet
R-301
title
Research
subtitle
Published work and experimental directions
discipline
R · Research
scale
1:1
revision
B
issued
2026.09
refs
A-101, W-400
msc
Lviv Polytechnic National University, 2021–2023, distinction
btech
RGPV Bhopal, 2017–2021, honours
isbn
978-3-031-60815-5

S-201Structural

Structural Principles

The decisions that outlast implementation

System design, infrastructure patterns, and how I think about structure.

Scale
1:20
Rev
C
Issued
2026.09
1Section through a governed platformS-201 · 1:20

Architecture is the set of decisions that outlast implementation. In AI systems that means naming invariants early: the clinician or operator is always the final authority, every model output is traceable to a version and data slice, and ambiguous database questions never execute without preview, disambiguation, and explicit human confirmation. Those are not compliance checkboxes. They are structural choices that keep NL2SQL, RAG, and agent tool calls from becoming silent liabilities.

Unstructured EHR · as receivedPatient John D. 45y male arrived 09:30 complaining of sharp chest pain radiating to left arm. BP is 145/90, HR 98. Meds: lisinopril 10mg daily. No known allergies. Doctor notes slight diaphoresis. Sent for stat ECG and troponin levels.… parsing string
… regex failed
… layout parser confidence: 12%
Extracted intelligence · JSON
{
  "patient_age": 45,
  "patient_gender": "M",
  "symptoms": ["chest pain", "diaphoresis"],
  "vitals": {
    "bp": "145/90",
    "hr": 98
  },
  "medications": ["lisinopril 10mg"],
  "orders": ["ECG", "troponin"],
  "confidence_score": 0.98
}
Field
Conf.
patient_age
1.00
patient_gender
0.99
symptoms
0.97
vitals
0.99
medications
0.98
orders
0.98
… schema validated
… 6 of 6 fields resolved
… layout parser confidence: 98%
Fig. 1Drag the divider. The same clinical note, before and after the extraction pipeline. The point of the architecture is that the right-hand state is traceable back to the left.

For conversational and agentic platforms I default to a gateway-first layout: tiered routing (exact match → classifier → composed retrieval), hybrid dense-and-sparse search with cross-encoder reranking, per-tenant keys and caching, OpenRouter or Bedrock-backed model selection with fallbacks, and microservices bounded by failure domain: orchestration, retrieval, execution against real rows, guardrails with escalation flags, observability on every hop. Monorepos like AURIXA and forward-deployed stacks on AWS CDK (Lambda, API Gateway, RDS, Redis, secrets) are different packaging of the same idea: scale the concern that hurts, not the whole binary.

OBSERVABILITY ON EVERY HOPREQUESTGATEWAYper-tenant keys · cachingTIER 1 · EXACT MATCHTIER 2 · CLASSIFIERTIER 3 · COMPOSEDretrievalHYBRID SEARCHdense + sparse · rerankedMODEL SELECTIONwith fallbacksGUARDRAILSescalation flagsRESPONSE

A request the first tier can answer by exact match.

Gateway › Tier 1 · exact match › Guardrails

tiers tried: 1 of 3

The cheapest tier is asked first. Most of the routing is deciding how little work a request needs.

Fig. 2The gateway-first layout from the paragraph above. Choose a request; each routing tier is tried in turn and the first that can serve it does.

When throughput dominates (video ingest, restaurant demand sensing, finance onboarding), I reach for Kafka (or equivalent) ingestion, asynchronous inference, WebSocket fan-out for operators, and sub-200ms ingestion-to-decision budgets where the product promise requires it. Air-gapped and on-premise defense deployments add another axis: self-contained inference stacks without assuming a always-on cloud control plane.

MLOps is part of architecture, not an appendix: MLflow and W&B for experiment lineage, QLoRA when fine-tunes must be affordable, batching and quantization when p95 cost matters, Spark when batch feature work belongs off the request path, Terraform and CDK when environments must be reproducible. Local Docker parity with production is non-negotiable for the teams I lead: if staging cannot run the same contract as prod, UAT is theatre.

Drawing conventions

The same discipline applies to this site. It is issued as a drawing set with three states, and the controls that govern it are exposed rather than hidden.

Control scheduleDraft

These are the real controls for this sheet, not a specimen panel. Everything below is wired to the drawing you are looking at.

Recurring themes across engagements: minimal surface area, clear boundaries, performance as a requirement from day one, and the conviction that the best dependency is the one you do not add. Case studies with tradeoffs and metrics are in Work; longer-form philosophy in AI infrastructure philosophy and Essays.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleStructural PrinciplesRevCIssued2026.09S-201

S-201record

sheet
S-201
title
Structural Principles
subtitle
The decisions that outlast implementation
discipline
S · Structural
scale
1:20
revision
C
issued
2026.09
refs
W-400, B-500, E-600
invariants
operator authority · output traceability · no silent execution
default
gateway-first, tiered routing

W-401Works

Forward-Deployed: Multi-Tenant Fertility & Surrogacy AI

Three coordinated chatbots, a governed NL-to-SQL reporting agent, and a shared LLM platform layer across multiple client programs, with a roadmap to production on AWS CDK.

Scale
1:1
Rev
A
Issued
May 1, 2026
Reading
3 min
Engagement
Forward deployment · health platform
  • Python
  • FastAPI
  • React
  • Vite
  • AWS CDK
  • Lambda
  • RDS
  • Redis
  • OpenRouter
  • Playwright
Schedule of outcomes4 items
  1. 013coordinated AI surfaces
  2. 028000+validated schema columns
  3. 03100+legacy prompts retired
  4. 04Human-in-the-loop NL2SQL
ADMINSCHEMA ROUTING8000+ approved columnsINTENT CLASSIFICATIONDISAMBIGUATIONunder-specified joinsQUERY PREVIEWHUMAN CONFIRMATIONmandatoryAUDITED EXECUTIONRESULT

A question whose tables and joins are unambiguous.

Schema routing › Intent classification › Query preview › Human confirmation › Audited execution

previewed before runningexecuted after confirmation

Even the easy case stops at a person. The confirm step is the product, not friction on top of it.

Fig. 1The governed NL-to-SQL agent. Choose a question; whatever path it takes, nothing runs until a person confirms it.

Problem

Fertility and surrogacy operations run on sensitive intake, deep domain knowledge, and reporting that must survive audit, not on a single generic chatbot. Product needed three coordinated experiences: intake assistance, knowledge support for staff, and administrative reporting. Executives needed confidence that database-facing AI would not hallucinate joins or execute destructive SQL. Engineering needed one authentication model, shared APIs, and release governance across multiple client brands.

Role & delivery model

Forward-deployed as lead full-stack AI engineer, I owned architecture through UAT: workbook reviews with product, Jira and Slack decision logs, SQL reviews with product owners, and evidence-backed scope negotiation when stakeholders disagreed. I also extended a personal real-estate AI prototype into a production SaaS platform for a US federal use case: prototype to deployed product across workflows, application architecture, and infrastructure.

Platform architecture

Shared LLM layer for all client bots:

  • Central gateway with tiered routing: exact match → classifier → composed retrieval
  • Hybrid dense + sparse retrieval with cross-encoder reranking
  • OpenRouter-based model selection, per-bot keys, caching, fallback handling for cost- and latency-aware inference

Application stack: FastAPI microservices, React/Vite administration UX, JWT-based multi-bot routing, Docker local parity with AWS CDK-managed Lambda, API Gateway, RDS, Redis, secrets, and storage. Regression coverage via Playwright and pytest.

Multi-tenant product surface: Intake, knowledge, and admin reporting bots sharing auth and integration patterns, not three forked codebases.

Governed NL-to-SQL reporting agent

The hardest wedge was schema ambiguity: complex multi-join questions where a wrong table choice looks plausible. Architecture:

  1. Schema-grounded routing and intent classification
  2. Query preview and disambiguation when joins are under-specified
  3. Mandatory human confirmation before execution
  4. Audit-friendly execution path suitable for C-suite and product readouts

I presented this design to executive stakeholders and aligned teams on risk, scope, timeline, and UAT acceptance criteria.

Schema corpus reset

With client product we retired a 100+ prompt legacy catalog in favor of a stakeholder-validated reporting corpus: roughly 8000+ approved columns across 193 operational tables. Deliverables included manifest SQL, LLM intent routing, schema-coverage tooling, and staging sign-off harnesses that achieved full automated pass on the validated batch.

Lessons

  1. Human confirmation is a feature, not friction: for NL2SQL in regulated domains, the confirm step is the product.
  2. One gateway beats N bespoke bots: per-bot keys and routing tiers share cost controls and observability.
  3. Forward deployment is translation: the same technical decision must read as risk reduction for executives and as a sprint plan for engineers.
  4. Schema work is product work: eight thousand columns of agreement is what makes the agent trustworthy.
Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleForward-Deployed: Multi-Tenant Fertility & Surrogacy AIRevAIssuedMay 1, 2026W-401

W-401record

sheet
W-401
title
Forward-Deployed: Multi-Tenant Fertility & Surrogacy AI
subtitle
Three coordinated chatbots, a governed NL-to-SQL reporting agent, and a shared LLM platform layer across multiple client programs, with a roadmap to production on AWS CDK.
discipline
W · Works
scale
1:1
revision
A
issued
May 1, 2026
refs
none
series
Works
words
427
stack
Python, FastAPI, React, Vite, AWS CDK, Lambda, RDS, Redis, OpenRouter, Playwright
metrics
3 coordinated AI surfaces · 8000+ validated schema columns · 100+ legacy prompts retired · Human-in-the-loop NL2SQL

W-402Works

AIDA: Health Programme Intelligence You Can Brief From

Turning facility assessment rows into screening rates, management gaps, and district rollups, with optional LLM narratives that never invent the numbers.

Scale
1:1
Rev
A
Issued
Apr 20, 2026
Reading
3 min
Engagement
Open source · ax5hay/AIDA
  • NestJS
  • Next.js
  • Prisma
  • PostgreSQL
  • TypeScript
  • npm workspaces
Schedule of outcomes3 items
  1. 01Deterministic KPI engine
  2. 02Parity ANC workspace in same monorepo
  3. 03Live demo at aida-demo.merakiel.in

Problem

State and district health programmes collect monthly assessment data per facility: who was identified vs managed, ANC screening counts, delivery signals, neonatal follow-up. The raw tables are unusable in a briefing unless you can defend rates, gaps (identified minus managed on the same condition keys), and time-bounded views filtered by period, district, and facility.

Spreadsheets break at scale. Dashboards that let the UI compute rates diverge from the API. Narrative “AI insights” that hallucinate counts destroy trust with programme officers.

Architecture

AIDA enforces a strict boundary: the UI never opens a database connection.

THE UI NEVER OPENS A DATABASE CONNECTIONROWSPOSTGRESQLmonthly assessmentsPRISMA@aida/dbNEST APIANALYTICS-ENGINErates · validationML-ENGINEcorrelations · z-scoresAI-ENGINEoptional narrativesNEXT.JSBRIEFING FIGURE

HIV tested, as a share of everyone registered for ANC, for one district and period.

PostgreSQL › Prisma › Nest API › analytics-engine › Next.js

deterministicmodel involved: no

Policy-grade math. The rate is defined once, in the analytics engine, and nowhere else.

Fig. 1Choose a request. Whichever engine answers it, the numbers are computed behind the API and the web app only displays them.

The same monorepo ships Parity: an ANC capture, analytics, and observation workspace (parity-web, parity-api, @aida/parity-core). A product hub links the two apps via configurable PARITY_WEB_URL / AIDA_WEB_URL, so one deployment can expose both surfaces without rebuilding for new demo hosts.

Package Responsibility
analytics-engine Mortality, LBW, preterm, institutional delivery mix, screening rates, management gaps
ml-engine Pearson correlations, z-score anomaly flags on delivery metrics
ai-engine POST /v1/ai/insights: narrates a JSON snapshot the API already computed
parity-core ANC indicator schema, validation, analytics bundle

Tech stack

  • API: NestJS modules covering analytics, metrics, facilities, ingestion, optional AI
  • Web: Next.js App Router covering overview, analytics suite, explorer, correlations, help
  • Data: Prisma schema as source of truth; identified vs managed in parallel section tables
  • Deploy: Docker Compose for Postgres + API + web; demo-start.sh for full Parity + AIDA stack
  • Demo: aida-demo.merakiel.in

Query params (from, to, district, facilityId) slice every analytics endpoint consistently, giving shareable URLs for filtered views.

Tradeoffs

  • Deterministic KPIs vs. ML flourishes: Correlations and anomalies are clearly separated from headline rates so officers know what is policy-grade math vs exploratory stats.
  • Optional LLM: The product is fully usable with no model server. Narratives only activate when AI_INSIGHTS_ENABLED and an OpenAI-compatible endpoint are configured, and the model receives counts, not raw PHI lists.
  • Monorepo complexity: Two APIs and two web apps share one database package. The cost is wiring CORS (WEB_ORIGIN, PARITY_WEB_ORIGIN) correctly; the win is one schema and one analytics definition.

Metrics & capabilities

  • Screening coverage: e.g. hiv_tested ÷ summed total_anc_registered via screeningRates in the analytics engine.
  • Management gaps: Identified vs managed section tables with validation that managed ≤ identified where schema implies it.
  • API surface: Overview KPIs, district rollup, clinical cross-section, assessment explorer, ingestion with server-side validation.
  • Performance: Gzip on JSON, 30s in-memory cache on hot analytics paths, narrow Prisma selects, TanStack Query with keepPreviousData on filter changes.

Lessons

  1. Keep derived definitions in one engine. Rate math duplicated in SQL and React is how programmes lose faith in dashboards.
  2. Thin UI, fat API. The web app mirrors query params; it does not re-derive epidemiology.
  3. LLM as narrator, not author. Insights POST sends the same JSON shape as /analytics/overview: the model explains, it does not count.
  4. Parity extends capture, not replacement. ANC field discipline and AIDA rollups address different moments in the same programme workflow.

Source

Setup, API routes, and Parity docs: github.com/ax5hay/AIDA

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAIDA: Health Programme Intelligence You Can Brief FromRevAIssuedApr 20, 2026W-402

W-402record

sheet
W-402
title
AIDA: Health Programme Intelligence You Can Brief From
subtitle
Turning facility assessment rows into screening rates, management gaps, and district rollups, with optional LLM narratives that never invent the numbers.
discipline
W · Works
scale
1:1
revision
A
issued
Apr 20, 2026
refs
none
series
Works
words
546
stack
NestJS, Next.js, Prisma, PostgreSQL, TypeScript, npm workspaces
metrics
Deterministic KPI engine · Parity ANC workspace in same monorepo · Live demo at aida-demo.merakiel.in

W-403Works

Agentic Video Intelligence for 24/7 Operations

Real-time CCTV triage with LangGraph and MCP: Kafka ingestion, sub-200ms paths, and defense-grade on-prem deployment.

Scale
1:1
Rev
A
Issued
Apr 1, 2026
Reading
1 min
Engagement
Defense-adjacent · video intelligence
  • LangGraph
  • MCP
  • FastAPI
  • Kafka
  • WebSockets
  • React
  • Next.js
  • Python
Schedule of outcomes3 items
  1. 01~70%less analyst intervention
  2. 02Sub-200msingestion-to-decision
  3. 03On-prem defense deployments
LANGGRAPH + MCP · PERSISTENT MEMORY AND TOOL USE ACROSS STEPSFEEDKAFKA INGESTIONasynchronous inferenceDETECTION AGENTEVENT CLASSIFICATIONWEBSOCKET STREAMlive dashboardALERT ESCALATIONoperational playbooksNO ALERT

An ordinary frame on an ordinary feed.

Kafka ingestion › Detection agent › Event classification

no analyst involved

This is most of the footage, and it is the part that used to burn analysts out. It is classified and let go.

Fig. 1Choose what a camera sees. Triage is a graph with state, so a routine frame and an incident leave by different doors.

Problem

Security and defense operators cannot watch every feed. Analysts burn out on false positives; true incidents arrive late because triage is manual. The platform targets real-time video intelligence: detect, classify, escalate, and drive downstream workflows without requiring a human on every frame.

Architecture

Stateful multi-agent orchestration with LangGraph and MCP:

  • Detection and event classification agents with persistent memory and tool use
  • Alert escalation pipelines that respect operational playbooks
  • FastAPI microservices behind WebSocket event streams for live dashboards

Ingestion: Kafka-based pipelines with asynchronous inference, engineered for sub-200ms ingestion-to-decision latency on hot paths.

Deployment modes: Cloud-native for iteration; self-contained inference stacks for air-gapped, on-premise defense infrastructure where outbound cloud calls are not an option.

Outcomes

  • Analyst intervention reduced by approximately 70% through automated detection triage and workflow handoff
  • End-to-end ownership of defense-sector deployments: infrastructure, inference, and React/Next.js operational dashboards

Lessons

  1. Agents need state, not just prompts: classification and escalation are graphs, not single-shot completions.
  2. Latency is a trust metric: operators abandon dashboards that lag the wall of cameras.
  3. Design for disconnected environments early: packaging models and brokers for on-prem avoids a rewrite when classification moves to classified networks.
Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAgentic Video Intelligence for 24/7 OperationsRevAIssuedApr 1, 2026W-403

W-403record

sheet
W-403
title
Agentic Video Intelligence for 24/7 Operations
subtitle
Real-time CCTV triage with LangGraph and MCP: Kafka ingestion, sub-200ms paths, and defense-grade on-prem deployment.
discipline
W · Works
scale
1:1
revision
A
issued
Apr 1, 2026
refs
none
series
Works
words
199
stack
LangGraph, MCP, FastAPI, Kafka, WebSockets, React, Next.js, Python
metrics
~70% less analyst intervention · Sub-200ms ingestion-to-decision · On-prem defense deployments

W-404Works

Neural Map Personal Site: Spatial Navigation as Homepage

A scroll-driven Three.js constellation where clusters replace nav bars, with hover, fly-in, and iframe overlays for essays, work, and writing.

Scale
1:1
Rev
A
Issued
Feb 28, 2026
Reading
3 min
Engagement
Open source · ax5hay/akshaybajpai.com
  • Next.js 15
  • Three.js
  • TypeScript
  • remark
  • GitHub Pages
Schedule of outcomes3 items
  1. 015interactive clusters
  2. 02Static export to out/
  3. 03p99-friendly LineSegments renderer

Problem

Personal sites default to the same pattern: hero, grid of cards, footer links. The content is fine; the metaphor is tired. This site asks: what if the homepage were a space you move through, where sections are places, not menu items?

Constraints stayed strict: free hosting (GitHub Pages), static export, readable labels, no laggy WebGL. The map had to feel alive without becoming an unreadable shader demo.

Architecture

The homepage is one scene, not a stack of marketing sections.

READERSCROLL CAMERAz tied to scrollYRAYCAST HOVERCLUSTER NUCLEUS5 click targetsRAW MODEdev snippets as labelsLABEL AND DIMnode scales, edges litFLY-TO, THEN OVERLAYclosed → flying → open

The reader scrolls the page.

Scroll camera

camera advances into the network

There were no sections to scroll past. Scrolling moved you through one scene.

Fig. 1The previous version of this site, as its interaction model. Choose an input and see what the scene did with it.

Five clusters in a ring (Essays, About, Work, Blog, Contact), each with a nucleus (click target) and orbiting thought-nodes. InstancedMesh for nodes; LineSegments with in-place buffer updates for edges (no per-frame geometry allocation). Cross-cluster lines appear only when nodes from different clusters drift close, keeping the graph legible.

Inner pages use a conventional PageShell: header, placard layout, markdown content collections. The neural map is the front door; everything else is a room you enter through it.

Tech stack

Layer Choice
Framework Next.js 15 App Router, output: 'export', trailingSlash: true
3D Three.js: lazy-loaded via dynamic(..., { ssr: false })
Content Markdown in content/: blog, essays, work; remark + gray-matter
SEO Per-route metadata, JSON-LD, sitemap, post-build RSS
Fonts Instrument Serif, IBM Plex Sans/Mono via next/font
Deploy GitHub Actions → out/ + CNAME → www.akshaybajpai.com

Reduced motion: If prefers-reduced-motion, the scene skips orbit/drift but interaction and overlays still work.

Tradeoffs

  • Iframe overlays vs. SPA transitions: Sections load in a native fullscreen panel: simple, works with static export, avoids routing the entire site through WebGL.
  • LineSegments vs. tubes/shaders: An earlier tube-and-shader experiment looked striking but tanked frame times. The current renderer prioritizes stable 60fps on consumer hardware.
  • No header on hero: Navigation lives in the map; scroll reveals hints, not a nav bar. Inner pages restore standard chrome.

Metrics & capabilities

  • Clusters: 5 nuclei, 80 instanced nodes, up to 900 curved edges with 6 samples per segment.
  • Interaction: Raycast hover, fly-to animation on select, overlay phases closed → flying → open.
  • Build: ~29 static routes; homepage JS kept small by code-splitting Three.js.
  • CI: Build verifies out/index.html and writes CNAME; deploy-pages with retry logic.

Lessons

  1. Performance is a design constraint. The beautiful version and the fast version had to be the same version.
  2. Labels must never hide. Hover text and hints exist because mystery navigation is not immersive; it is hostile.
  3. Static export shapes the UX. Iframe overlays and client-only Three.js are consequences of GitHub Pages, not accidents.
  4. The map is navigation, not decoration. Every cluster maps to real content; raw mode is an Easter egg for builders, not the primary UI.

Source

Live: www.akshaybajpai.com · Code: github.com/ax5hay/akshaybajpai.com

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleNeural Map Personal Site: Spatial Navigation as HomepageRevAIssuedFeb 28, 2026W-404

W-404record

sheet
W-404
title
Neural Map Personal Site: Spatial Navigation as Homepage
subtitle
A scroll-driven Three.js constellation where clusters replace nav bars, with hover, fly-in, and iframe overlays for essays, work, and writing.
discipline
W · Works
scale
1:1
revision
A
issued
Feb 28, 2026
refs
none
series
Works
words
525
stack
Next.js 15, Three.js, TypeScript, remark, GitHub Pages
metrics
5 interactive clusters · Static export to out/ · p99-friendly LineSegments renderer

W-405Works

AURIXA: Conversational AI Orchestration at Scale

Multi-tenant microservices platform for real-time LLM routing, RAG, agent execution, safety guardrails, and voice, built as a production Turborepo monorepo.

Scale
1:1
Rev
A
Issued
Feb 18, 2026
Reading
4 min
Engagement
Open source · ax5hay/AURIXA
  • TypeScript
  • Python
  • FastAPI
  • Fastify
  • Next.js 15
  • PostgreSQL
  • Redis
  • Docker
  • Turborepo
Schedule of outcomes3 items
  1. 018Python services + API gateway
  2. 02Pluggable OpenAI / Claude / Gemini / local LLMs
  3. 03E2E pipeline with emergency escalation

Problem

Most “chat with your data” demos collapse under real constraints: multiple tenants, provider failover, cost-aware routing, tool execution against real databases, safety checks before a response ships, and observability when eight services are in the path. AURIXA exists to answer a single question: how do you run conversational AI like infrastructure, not like a script?

The platform targets healthcare-adjacent workflows (appointments, insurance checks, prescription refills) but the architecture is domain-agnostic: stateless services, async orchestration, and a gateway that can scale each concern independently.

Architecture

Every request enters through a Fastify API Gateway (rate limits, CORS, WebSocket proxy, structured logging), then flows into an Orchestration Engine that coordinates the pipeline:

OBSERVABILITY CORE :8008 · TELEMETRY FROM EVERY HOPUSERGATEWAY:3000ORCHESTRATION:8001LLM ROUTER:8002AGENT RUNTIME:8003EXECUTION ENGINE:8007RAG SERVICE:8004GUARDRAILS:8005RESPONSE

“Can I see someone on Thursday?”

Gateway :3000 › Orchestration :8001 › LLM Router :8002 › Agent Runtime :8003 › Execution Engine :8007 › Guardrails :8005

6 servicesrequires_escalation: false

Agent branch. The execution engine reads and writes tenant-scoped rows, so the answer is a real slot.

Fig. 1Choose a request. The path it takes is traced through the services and ports this section names.

Execution Engine actions are not mocks; they read and write tenant-scoped records: get_appointments, create_appointment, check_insurance, get_availability, request_prescription_refill. Safety Guardrails flag phrases like chest pain or stroke and set requires_escalation rather than pretending the model is a clinician.

Frontends ship in the same monorepo: unified admin dashboard (playground, tenants, service health), patient portal, and hospital portal, all Next.js 15.

Tech stack

Layer Choices
Monorepo Turborepo + pnpm workspaces
Gateway Fastify 5, TypeScript
Services FastAPI (Python 3.11+): orchestration, LLM router, RAG, agents, safety, voice, execution, observability
Data PostgreSQL 16, Redis 7
LLM layer Shared llm-clients package: OpenAI, Anthropic, Gemini, local models
Infra Docker Compose locally; K8s + Terraform templates for AWS
UI Next.js 15 dashboards, shared ui-kit

Response caching (TTL 300s) and telemetry emission on orchestration, routing, and RAG reduce cost and make the playground’s “Run All Tests” panel meaningful.

Tradeoffs

  • Microservices vs. velocity: Eight services plus three frontends is heavy for a solo builder, but it mirrors how production AI platforms actually fail: at routing, safety, and observability boundaries. The split buys independent deploy and clear ownership per concern.
  • Healthcare demo data vs. generic core: Sample patients and appointments anchor the execution engine, yet the gateway–orchestration–router pattern transfers to any vertical with tool calls and RAG.
  • Python + TypeScript split: Gateway and auth stay in Node; ML-heavy paths stay in FastAPI. Two runtimes, one contract: Pydantic schemas and shared auth utilities.

Metrics & capabilities

  • Services: API Gateway + 8 FastAPI microservices, each with health endpoints and playground coverage.
  • Pipeline: Intent classification → RAG or agent branch → generation → safety validation in one orchestrated pass.
  • Multi-tenant: Admin API for tenant and patient creation; knowledge articles scoped per tenant for RAG.
  • Ops: Playground dashboard runs full E2E tests, surfaces per-service latency, and visualizes Intent → RAG/Agent → Generate → Safety steps.

Lessons

  1. Orchestration is the product. Routing and RAG are commodities; the engine that sequences them, caches, and escalates is what makes the system trustworthy.
  2. Safety belongs in the graph, not in the prompt. A dedicated guardrails service with explicit escalation flags beats hoping the LLM self-censors.
  3. DB-backed tools ground agents. Appointment and insurance actions against real rows prevent the “helpful but fictional” failure mode.
  4. Observability from service one. When eight hops are normal, per-service metrics and audit logs are not optional.

Source

Full architecture, service ports, and setup: github.com/ax5hay/AURIXA

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAURIXA: Conversational AI Orchestration at ScaleRevAIssuedFeb 18, 2026W-405

W-405record

sheet
W-405
title
AURIXA: Conversational AI Orchestration at Scale
subtitle
Multi-tenant microservices platform for real-time LLM routing, RAG, agent execution, safety guardrails, and voice, built as a production Turborepo monorepo.
discipline
W · Works
scale
1:1
revision
A
issued
Feb 18, 2026
refs
none
series
Works
words
618
stack
TypeScript, Python, FastAPI, Fastify, Next.js 15, PostgreSQL, Redis, Docker, Turborepo
metrics
8 Python services + API gateway · Pluggable OpenAI / Claude / Gemini / local LLMs · E2E pipeline with emergency escalation

W-406Works

Alzheimer's Classification: Master's Thesis Research

Benchmarking nine machine learning models on OASIS longitudinal biomarkers for dementia detection: rigorous preprocessing, grid search, and clinical interpretability.

Scale
1:1
Rev
A
Issued
Jan 25, 2026
Reading
3 min
Engagement
Master's thesis · ax5hay/AlzheimersDiagnosis
  • Python
  • scikit-learn
  • pandas
  • Jupyter
  • OASIS dataset
Schedule of outcomes3 items
  1. 019models compared
  2. 025-foldcross-validation
  3. 03ROC-AUC + recall for clinical sensitivity
OASISFIRST VISIT ONLY150 subjectsENCODEbinary gender + labelMISSING SESDROP ROWSstrategy A · 8 rowsMEDIAN IMPUTEstrategy B · by EDUCSCALE AND SPLIT75/25 · 5-fold CV9 MODELS

Strategy A: drop the rows where SES is missing.

First visit only › Encode › Missing SES › Drop rows › Scale and split

8 rows removedrandom_state=0

Nothing is invented, at the cost of a smaller sample from an already small cohort.

Fig. 1The thesis pipeline. Choose how missing socioeconomic status is handled; it is a scientific choice, and it changes the result.

Problem

Early cognitive decline is easy to miss in routine care. The research question: can a small set of clinical and neuroimaging biomarkers reliably separate demented from non-demented patients in a cross-sectional snapshot, and which algorithms balance accuracy with recall for disease detection?

This was not a production deployment; it was thesis-grade methodology: explicit preprocessing choices, imputation strategies compared side by side, and nine models tuned through grid search with held-out test evaluation.

Dataset & features

Source: OASIS Longitudinal Dataset, 150 subjects at first clinical visit, 8 predictive variables.

Category Features
Clinical Gender, age, years of education (EDUC), socioeconomic status (SES)
Cognitive MMSE (Mini-Mental State Examination)
Neuroimaging Estimated intracranial volume (eTIV), normalized whole brain volume (nWBV), atlas scaling factor (ASF)

MMSE and brain volume metrics showed the strongest separation between groups: MMSE ranges clustered around 25–30 for non-demented vs 17–30 for demented cohorts.

Methodology

  1. Selection: First-visit rows only for cross-sectional analysis.
  2. Encoding: Binary gender; dementia label standardized to binary target.
  3. Missing values: Strategy A drops rows with missing SES (8 rows). Strategy B uses EDUC-stratified median imputation.
  4. Scaling: MinMax normalization on training folds.
  5. Split: 75% train/validation (5-fold CV), 25% held-out test; random_state=0 for reproducibility.

Models evaluated

Nine distinct approaches with grid-search hyperparameters:

  • Logistic Regression (with and without imputation)
  • SVM with RBF, linear, polynomial, and sigmoid kernels
  • Decision Tree (max depth search)
  • Random Forest (estimators, features, depth)
  • AdaBoost (estimators, learning rate)

Metrics: Accuracy, recall (dementia sensitivity), ROC-AUC, confusion matrices. Feature importance from tree-based models; Graphviz export of optimal decision tree structure.

Results & insights

  • Class imbalance required emphasizing recall, not accuracy alone; missing dementia is costlier than a false alarm in screening context.
  • MMSE dominated feature importance across tree ensembles; neuroimaging ratios added signal but cognitive score carried most discriminative power.
  • SVM and ensemble methods (Random Forest, AdaBoost) competed on AUC; logistic regression anchored interpretability.
  • Imputation strategy materially shifted performance; documenting both paths was essential for thesis rigor.

Lessons

  1. Preprocessing is the experiment. Imputation vs complete-case analysis is a scientific choice, not a footnote.
  2. Recall is the clinical metric. Optimize for the error you cannot afford.
  3. Interpretability has value. Decision trees and feature importance charts support clinician conversation even when ensembles win on AUC.
  4. Reproducibility is non-negotiable. Fixed random seeds and explicit train/test walls keep thesis results defensible.

Source

Notebook, methodology, and dataset notes: github.com/ax5hay/AlzheimersDiagnosis

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAlzheimer's Classification: Master's Thesis ResearchRevAIssuedJan 25, 2026W-406

W-406record

sheet
W-406
title
Alzheimer's Classification: Master's Thesis Research
subtitle
Benchmarking nine machine learning models on OASIS longitudinal biomarkers for dementia detection: rigorous preprocessing, grid search, and clinical interpretability.
discipline
W · Works
scale
1:1
revision
A
issued
Jan 25, 2026
refs
none
series
Works
words
422
stack
Python, scikit-learn, pandas, Jupyter, OASIS dataset
metrics
9 models compared · 5-fold cross-validation · ROC-AUC + recall for clinical sensitivity

W-407Works

Agentic Finance & Multimodal Healthcare AI

LangGraph workflows across portfolio logic and onboarding, plus QLoRA clinical imaging pipelines on HIPAA-aware AWS with hybrid RAG at sub-800ms p95.

Scale
1:1
Rev
A
Issued
Nov 1, 2025
Reading
2 min
Engagement
Consulting · finance & health
  • LangGraph
  • QLoRA
  • PyTorch
  • AWS
  • Python
  • RAG
  • Hugging Face
Schedule of outcomes4 items
  1. 01~60%fewer manual touchpoints
  2. 0294%precision on diagnostic imaging
  3. 0380%inference cost reduction
  4. 04Sub-800msp95 RAG
CLIENTLANGGRAPH ROUTERstateful tool routingRISK PROFILING APIPORTFOLIO REBALANCINGONBOARDING UTILITIESCROSS-STEP MEMORYkept across stepsNEXT STEP

A new client starts onboarding.

LangGraph router › Onboarding utilities › Cross-step memory

retries and fallbacks on tool failure

What onboarding collects is written to memory, so the client is not asked for it again two steps later.

Fig. 1The finance graph. Choose a step; the router picks the tool, and what it learns is kept for the steps that follow.

Problem

Wealth and health domains both punish “helpful” hallucinations. Finance workflows need stateful tool routing across risk APIs, rebalancing logic, and onboarding, with retries, fallbacks, and memory that survives multi-step conversations. Healthcare imaging pipelines need precision and compliance: fine-tuned models inside HIPAA/GDPR-aware infrastructure, not notebook accuracy.

Finance: LangGraph agentic workflows

Built stateful multi-agent orchestration:

  • Tool routing across risk profiling APIs, portfolio rebalancing, and onboarding utilities
  • Fallback handling, retry logic, and cross-step memory persistence
  • Approximately 60% reduction in manual touchpoints for supported journeys

Healthcare: multimodal pipelines

  • EHR and diagnostic image analysis with fine-tuned open-source LLMs via QLoRA
  • 94% precision on the targeted imaging tasks within compliant AWS boundaries
  • 80% inference cost reduction through quantization and batching optimizations

Production RAG

Hybrid retrieval, contextual reranking, and domain guardrails sustaining sub-800ms p95 latency: the bar where operators treat the system as interactive, not batch.

Lessons

  1. Memory and routing are the finance product: the base model is interchangeable; the graph is not.
  2. Cost is an architecture input: QLoRA and batching decisions belong beside latency SLOs.
  3. Guardrails beat bigger models: domain constraints on retrieval and generation outperform raw parameter count for compliance-sensitive text.
Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAgentic Finance & Multimodal Healthcare AIRevAIssuedNov 1, 2025W-407

W-407record

sheet
W-407
title
Agentic Finance & Multimodal Healthcare AI
subtitle
LangGraph workflows across portfolio logic and onboarding, plus QLoRA clinical imaging pipelines on HIPAA-aware AWS with hybrid RAG at sub-800ms p95.
discipline
W · Works
scale
1:1
revision
A
issued
Nov 1, 2025
refs
none
series
Works
words
201
stack
LangGraph, QLoRA, PyTorch, AWS, Python, RAG, Hugging Face
metrics
~60% fewer manual touchpoints · 94% precision on diagnostic imaging · 80% inference cost reduction · Sub-800ms p95 RAG

W-408Works

Publishing AI Automation: Multi-Channel Support at Scale

Multi-channel author query automation with intent classification, hybrid RAG, unified identity across WhatsApp and email, and confidence-gated human escalation.

Scale
1:1
Rev
A
Issued
Sep 17, 2025
Reading
3 min
Engagement
Publishing · client engagement
  • Node.js
  • TypeScript
  • Express
  • Supabase
  • OpenAI GPT-4
  • Redis
  • Docker
Schedule of outcomes3 items
  1. 015communication channels
  2. 02Hybrid semantic + keyword RAG
  3. 03Entity extraction for ISBNs and titles

Problem

A publishing house fields the same questions across WhatsApp, Instagram, email, web, and SMS: order status, manuscript guidelines, ISBN lookups, royalty queries. Agents context-switch between channels; answers drift from the knowledge base; low-confidence guesses erode author trust.

The goal was not a chatbot widget. It was an automation platform: classify intent, retrieve the right policy paragraph, score confidence, escalate to humans when unsure, and keep one identity graph so “the same author” is recognized whether they DM or email.

Architecture

AUTHORQUERY PROCESSORintent · entitiesKNOWLEDGE LAYERsemantic + keywordIDENTITY UNIFICATIONfuzzy match, 5 channelsESCALATION QUEUERESPONSE GENERATORtemplates + GPT-4SENT · WHATSAPP

An author sends an ISBN and asks where their order is.

Query processor › Knowledge layer › Identity unification › Response generator

ISBN extracted from free textconfidence: above threshold

The formatter keeps it short for WhatsApp. The facts come from the knowledge base, not from the formatter.

Fig. 1Choose an incoming message. One pipeline serves all five channels, and a low-confidence answer goes to a person instead of the author.

Identity unification normalizes handles and fuzzy-matches profiles so a WhatsApp thread and an email thread can attach to one author record. Platform-specific formatters adapt the same factual answer to WhatsApp brevity vs email structure without maintaining five separate bots.

Tech stack

Component Implementation
API Express + TypeScript: queries, webhooks, health
Core queryProcessor.ts: pipeline orchestration
Knowledge ragSystem.ts: hybrid search over knowledge base documents
Database Supabase (PostgreSQL): profiles, conversations, audit
AI OpenAI GPT-4: intent classification, response drafting
Cache Redis: session and hot retrieval paths
Ops Docker Compose; optional Prometheus/Grafana hooks

Structured logging with request correlation IDs ties a webhook receipt to its RAG retrieval and final response for support debugging.

Tradeoffs

  • Custom code vs. no-code: Full control over confidence thresholds, Supabase schema, and channel formatters, at the cost of owning the integration layer instead of a SaaS bot builder.
  • GPT-4 for intent + generation: Higher quality for messy author language; mitigated by retrieval-first answers and escalation on low scores.
  • Supabase as backend: Fast iteration for a single-tenant publisher deployment; not multi-tenant SaaS out of the box without further isolation work.

Metrics & capabilities

  • Channels: WhatsApp, Instagram, Email, Web, SMS, with a unified identity layer across all five.
  • Entities: Extraction for ISBNs, emails, book titles from free-text queries.
  • RAG: Hybrid search using semantic embeddings plus keyword fallback for exact policy clauses.
  • Reliability: Health checks, graceful degradation, human escalation path for low-confidence classifications.

Lessons

  1. Confidence scores are a product feature. Authors prefer a human handoff over a wrong answer about royalties.
  2. Identity is the hidden integration cost. Channel-specific IDs must collapse to one profile or context is lost on every new message.
  3. Formatters ≠ models. Keep facts in one place; let the formatter adapt tone and length per channel.
  4. Observability beats prompt tweaking. Correlation IDs surfaced more failures than any single prompt revision.

Source

Repository and setup: github.com/ax5hay/ai-automation-main

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitlePublishing AI Automation: Multi-Channel Support at ScaleRevAIssuedSep 17, 2025W-408

W-408record

sheet
W-408
title
Publishing AI Automation: Multi-Channel Support at Scale
subtitle
Multi-channel author query automation with intent classification, hybrid RAG, unified identity across WhatsApp and email, and confidence-gated human escalation.
discipline
W · Works
scale
1:1
revision
A
issued
Sep 17, 2025
refs
none
series
Works
words
471
stack
Node.js, TypeScript, Express, Supabase, OpenAI GPT-4, Redis, Docker
metrics
5 communication channels · Hybrid semantic + keyword RAG · Entity extraction for ISBNs and titles

W-409Works

Restaurant SaaS Forecasting at Scale

Founding-engineer delivery for inventory and demand sensing: XGBoost and Prophet, Node APIs, Next.js ops dashboard, and 2M+ events/day pipelines.

Scale
1:1
Rev
A
Issued
Mar 1, 2025
Reading
1 min
Engagement
Restaurant SaaS · founding engineer
  • Node.js
  • Next.js
  • XGBoost
  • Prophet
  • Event-driven architecture
  • Python
Schedule of outcomes4 items
  1. 01200+restaurant pilot
  2. 0232%food waste reduction
  3. 032M+events/day
  4. 04Sub-100mspipeline latency
SITESEVENT PIPELINE2M+ events a dayDEMAND SENSINGNODE.JS SERVICESXGBOOST + PROPHETinventory forecastingORDER QUANTITIESprep and orderingOPS DASHBOARD

Demand across the group, as it moves.

Event pipeline › Demand sensing › Node.js services

sub-100ms on this path

The fast road. Holding that latency at this volume is a matter of partitioning and backpressure, not of the model.

Fig. 1Choose what an operator needs to know. The live view and the forecast share one event pipeline and part ways after it.

Problem

Restaurant groups run on thin margins and volatile demand. A SaaS pilot needed to prove forecasting-driven inventory and real-time demand sensing across hundreds of sites, not a dashboard demo, but operators trusting prep and order quantities daily.

What we built

As founding engineer, I led AI and full-stack delivery:

  • Forecasting: XGBoost and Prophet models for inventory optimization, validated across 200+ restaurants, with roughly 32% reduction in food waste in the pilot metrics we tracked
  • APIs & UX: Node.js services and a Next.js operator dashboard for franchise and central teams
  • Pipelines: Event-driven architecture load-tested at 2M+ events per day with sub-100ms latency on demand-sensing paths

Lessons

  1. Start with one workflow: prep and ordering beats “AI everywhere” on the menu.
  2. Event volume exposes design errors early: sub-100ms claims require honest partitioning and backpressure.
  3. Waste percentage is the executive metric: accuracy charts alone do not close restaurant pilots.

Broader industry framing: AI in Restaurant Automation.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleRestaurant SaaS Forecasting at ScaleRevAIssuedMar 1, 2025W-409

W-409record

sheet
W-409
title
Restaurant SaaS Forecasting at Scale
subtitle
Founding-engineer delivery for inventory and demand sensing: XGBoost and Prophet, Node APIs, Next.js ops dashboard, and 2M+ events/day pipelines.
discipline
W · Works
scale
1:1
revision
A
issued
Mar 1, 2025
refs
none
series
Works
words
165
stack
Node.js, Next.js, XGBoost, Prophet, Event-driven architecture, Python
metrics
200+ restaurant pilot · 32% food waste reduction · 2M+ events/day · Sub-100ms pipeline latency

W-410Works

Insurance Document Intelligence on AWS

Multimodal RAG, Claude and GPT-4 Vision, LayoutLMv3 extraction, and ensemble underwriting models: 95%+ accuracy at 10K+ queries/day.

Scale
1:1
Rev
A
Issued
Jun 1, 2024
Reading
2 min
Engagement
Insurance · document AI
  • AWS Bedrock
  • RAG
  • Claude
  • GPT-4 Vision
  • LayoutLMv3
  • LightGBM
  • Textract
  • Python
Schedule of outcomes3 items
  1. 0195%+extraction accuracy
  2. 0278%underwriting efficiency gain
  3. 0310K+queries/day
PDFTEXTRACT + TIKAOCR, high volumeLAYOUTLMV3layout-awareHYBRID RETRIEVALreranked · BedrockMULTIMODAL LLMClaude · GPT-4 VisionSTRUCTURED FEATURESBOOSTED ENSEMBLELightGBM, CatBoost, XGBGROUNDED ANSWER

A typed policy, and someone asking what it covers.

Textract + Tika › Hybrid retrieval › Multimodal LLM

10K+ queries a day on this path

The high-volume road. The answer is written from retrieved clauses, not from what a model remembers about insurance.

Fig. 1Choose a document and what is wanted from it. Language models extract and answer; gradient boosting decides.

Problem

Insurance operations ingest complex semi-structured documents: policies, endorsements, scans with tables and handwriting. Manual extraction does not scale; naive OCR misses layout; generic chatbots invent coverage details. The platform needed multimodal extraction, retrieval-grounded Q&A, and underwriting models that improve decisions without bypassing human sign-off.

Architecture

Document intelligence pipeline:

  • RAG with Claude 3.5 Sonnet and GPT-4 Vision for multimodal understanding
  • LayoutLMv3 for layout-aware extraction on challenging pages
  • OCR-to-RAG path using AWS Textract and Apache Tika for high-volume ingestion

Deployed on AWS Bedrock with hybrid retrieval, reranking, prompt optimization, and automated evaluation loops.

Downstream ML: LightGBM, CatBoost, and XGBoost ensembles that improved underwriting efficiency by ~78% on the workflows we targeted.

Scale & accuracy

  • 95%+ extraction accuracy across high-volume ingestion pipelines
  • 10K+ queries/day sustained on the RAG serving path

Lessons

  1. Layout is signal: vision + layout models beat text-only pipelines on insurance PDFs.
  2. Evaluation pipelines are production code: prompt tweaks without regression tests erode the 95% claim.
  3. Ensembles still win tabular underwriting: LLMs extract; gradient boosting decides when features are structured.

Architecture Design for Healthcare AI: shared themes on auditability and human override, applied across regulated domains.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleInsurance Document Intelligence on AWSRevAIssuedJun 1, 2024W-410

W-410record

sheet
W-410
title
Insurance Document Intelligence on AWS
subtitle
Multimodal RAG, Claude and GPT-4 Vision, LayoutLMv3 extraction, and ensemble underwriting models: 95%+ accuracy at 10K+ queries/day.
discipline
W · Works
scale
1:1
revision
A
issued
Jun 1, 2024
refs
none
series
Works
words
201
stack
AWS Bedrock, RAG, Claude, GPT-4 Vision, LayoutLMv3, LightGBM, Textract, Python
metrics
95%+ extraction accuracy · 78% underwriting efficiency gain · 10K+ queries/day

W-411Works

Clinical NLP from Unstructured EHRs

BioBERT and spaCy NER pipelines extracting 47+ structured variables at 91.5% accuracy, across 5K+ documents/day, with risk scoring at scale.

Scale
1:1
Rev
A
Issued
Jun 1, 2022
Reading
1 min
Engagement
Healthcare logistics · contract
  • BioBERT
  • spaCy
  • Python
  • scikit-learn
  • Clinical NLP
Schedule of outcomes4 items
  1. 0147+structured variables
  2. 0291.5%extraction accuracy
  3. 035K+documents/day
  4. 045K+patient records/day scoring
NOTECLINICAL NERBioBERT · spaCySTRUCTURED VARIABLES47+ per documentRISK SCORINGRF · boosting · SVMSTRUCTURED RECORD

A progress note, discharge summary or imaging report, abstracted into fields.

Clinical NER › Structured variables

91.5% extraction accuracy5K+ documents a day

What manual abstraction could do for a few charts a day, done for every document that arrives.

Fig. 1Choose how far a clinical narrative is taken: to structured variables, or on to a risk score.

Problem

Clinical research and operations teams sit on unstructured EHR narratives (progress notes, discharge summaries, imaging reports) while downstream analytics need structured variables and risk scores. Manual abstraction does not scale past a few charts per day.

Systems delivered

As Senior ML Engineer (contract), I built and productionized:

Clinical NLP extraction

  • 47+ structured variables from unstructured documents
  • 91.5% accuracy across pipelines processing 5K+ documents per day
  • BioBERT, spaCy NER, and custom sequence-labelling models tuned for clinical entities

Risk scoring

  • Supervised models (Random Forest, Gradient Boosting, SVM) for clinical risk scoring supporting triage
  • 5K+ patient records scored daily in production configuration

Lessons

  1. Entity lists are contracts: forty-seven variables only help if product defines each one defensibly.
  2. Domain embeddings matter: BioBERT-level priors beat general-language models on shorthand and abbreviations.
  3. Throughput is an NLP architecture problem: batching, model cascades, and fail-open paths keep 5K/day honest.

Springer chapter on ML for medical diagnosis (see Research) and Alzheimer's thesis work for the research side of the same clinical thread.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleClinical NLP from Unstructured EHRsRevAIssuedJun 1, 2022W-411

W-411record

sheet
W-411
title
Clinical NLP from Unstructured EHRs
subtitle
BioBERT and spaCy NER pipelines extracting 47+ structured variables at 91.5% accuracy, across 5K+ documents/day, with risk scoring at scale.
discipline
W · Works
scale
1:1
revision
A
issued
Jun 1, 2022
refs
none
series
Works
words
178
stack
BioBERT, spaCy, Python, scikit-learn, Clinical NLP
metrics
47+ structured variables · 91.5% extraction accuracy · 5K+ documents/day · 5K+ patient records/day scoring

E-601Essays

The Model That Refuses to Draw

Generative models predict pixels; JEPA predicts what the pixels mean. Why the refusal to reconstruct is the most important architectural decision in world models, and what it teaches anyone who builds systems.

Scale
1:1
Rev
A
Issued
Oct 6, 2026
Reading
7 min

There is a model family whose defining feature is a refusal. Asked to predict the future of a video, it will not draw the next frame. It predicts the representation of the next frame, a vector in a space it invented, and it is scored on whether that vector was right. The pixels never enter into it. This is the Joint Embedding Predictive Architecture, JEPA, which Yann LeCun set out in a 2022 position paper and which, four years on, is the backbone of the world-model work coming out of Meta, out of his new company, and out of a dozen labs following the same line.

The refusal looks like a limitation. It is the whole idea, and it carries a lesson that reaches well past machine learning.

What a frame costs

Every frame of video is almost entirely noise, in the sense that matters for prediction. The exact arrangement of leaves in a tree, the grain of a wall, the specular glint on a glass: none of it is predictable from the previous frame, and none of it needs to be. A model asked to reconstruct pixels spends most of its capacity on exactly this unpredictable texture, because that is where most of the loss lives. It learns to be a very good renderer of things that do not matter.

JEPA declines the bargain. An encoder maps the context (the frames you have) into a representation; a second encoder maps the target (the frames you are asked about) into the same space; a predictor is trained to get from one to the other. Whatever the encoders decide is not worth representing, the predictor is never penalised for missing. The model is free to throw the leaves away and keep the branch.

The engineering consequence is that the loss is measured in a space the model controls, which is also the obvious danger. If the encoders map everything to the same point, prediction is trivial and the model has learned nothing. This is representation collapse, and for years the field held it off with heuristics: a teacher network updated by moving average, a stop-gradient here, an asymmetric head there. They worked, and nobody could say exactly why.

Taking the heuristics out

In late 2025 Randall Balestriero and LeCun published LeJEPA, which replaces the heuristics with an argument. They show that the embeddings that minimise downstream prediction risk should follow an isotropic Gaussian, and they introduce a regulariser, SIGReg, that pushes the embedding distribution toward that shape using random one-dimensional projections, in linear time and memory. The predictive loss plus SIGReg is the whole method: one trade-off hyperparameter, no teacher and student, no stop-gradient, and a training loss that tracks the quality of the representation closely enough to be used for model selection without a labelled probe. They report stable training up to a 1.8-billion-parameter vision transformer, and an implementation in about fifty lines.

I find this the most interesting paper in the family, not for the result but for the shape of it. A working system held together by three unexplained tricks was replaced by a system held together by one explained constraint. That is what maturity looks like in any engineering discipline: the moment the folklore becomes a specification.

From understanding to doing

A representation is only worth having if something can act on it. V-JEPA 2, released by Meta in June 2025, was trained on more than a million hours of internet video and then, with about 62 hours of robot data, produced an action-conditioned predictor that could plan robot arm manipulation in a lab it had never seen: pick a goal image, imagine the latent consequences of candidate actions, pick the best, repeat. V-JEPA 2.1, in March 2026, made the representation denser (every token is supervised, visible and masked alike) and reported a twenty-point improvement in real-robot grasping success over its predecessor.

Notice what the robot never does. It never renders a frame of what the arm will do. It compares abstract summaries of possible futures and chooses between them. That is the claim of the whole programme in one sentence: planning is cheap in a space where only the relevant things are represented, and expensive in a space where everything is.

The other road

It would be dishonest to present this as the only live approach, because the most spectacular results of the last year came from models that do draw. DeepMind's Dreamer 4 trains an agent entirely inside a learned, generative simulator of Minecraft and is the first to reach diamonds from offline footage alone, a task of more than twenty thousand low-level actions, with a hundred times less action-labelled data than the previous offline agent. Genie 3 generates whole interactive worlds a person can walk through at twenty-four frames a second. These are generative world models, and they work.

The difference is in what the model is for. A simulator you want to look at, or train an agent inside with pixels as the interface, has to draw. A model whose only job is to let a controller choose between futures does not, and every pixel it draws is capacity taken from the judgement it exists to make. LeCun's bet, now funded at a billion dollars at AMI Labs, is that the second kind is what autonomy actually runs on. The year's evidence says both kinds are needed and that the field has not finished deciding where the boundary sits.

What it teaches the rest of us

I build systems for a living, most of them with language models in the loop, and I have come to read the JEPA refusal as a design principle rather than a research result.

Every system has a reconstruction loss it did not choose. A dashboard that re-derives rates in the browser is reconstructing what the API already computed. A retrieval pipeline that returns whole documents when the question needed one clause is predicting pixels. An agent that narrates every intermediate step to the user is rendering texture nobody asked for. In each case the capacity, human or machine, goes to reproducing the unpredictable surface of the thing rather than the part that determines what happens next.

The discipline JEPA imposes is to ask, before building anything, which representation the decision will actually be made in, and to refuse to compute anything below it. It is a harder question than it sounds, because the surface is what you can see and the representation is what you have to invent. But the systems I trust most, in production and in the literature, are the ones that answered it early and then declined, politely and permanently, to draw.

Sources

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleThe Model That Refuses to DrawRevAIssuedOct 6, 2026E-601

E-601record

sheet
E-601
title
The Model That Refuses to Draw
subtitle
Generative models predict pixels; JEPA predicts what the pixels mean. Why the refusal to reconstruct is the most important architectural decision in world models, and what it teaches anyone who builds systems.
discipline
E · Essays
scale
1:1
revision
A
issued
Oct 6, 2026
refs
none
series
Essays
words
1207

E-602Essays

Minimalism as Engineering

Why less is more in system design: fewer dependencies, fewer features, fewer moving parts, and how to get there without sacrificing capability.

Scale
1:1
Rev
A
Issued
Mar 5, 2025
Reading
3 min

Minimalism in engineering is not austerity for its own sake. It’s the recognition that every addition, every dependency, every feature, every configuration option, has a cost: in complexity, in failure modes, and in the cognitive load on the team. This essay is about when and how to choose less.

The cost of more

Every line of code can have a bug. Every dependency can break or change. Every feature can be misused or can interact badly with another feature. So the default should be: do we need this? If we add it, we own it forever (or until we remove it, which is also work). The cost of “more” is undercounted because we focus on the cost of building, not the cost of maintaining, debugging, and explaining. Minimalism is the habit of counting the full cost and adding only when the benefit clearly outweighs it.

Where minimalism pays off

Dependencies: The fewer you have, the less you have to upgrade, audit, and debug. Prefer the platform (browser APIs, standard library) and small, focused packages. When you add a dependency, abstract it behind your own interface so you can swap or remove it later.

Features: Every feature is a promise to support, document, and not break. Cut features that don’t earn their keep. Ship the smallest set that delivers the core value; add only when there’s evidence of need. “We might need it” is not evidence.

Configuration: Options multiply complexity. Every flag is a dimension in the matrix of “does it work?” Prefer convention over configuration. When you do add options, make the default the right choice for most users and document the rest.

Code: Delete dead code. Refactor to reduce surface area. A smaller codebase is easier to reason about, test, and change. Minimalism in code is not fewer lines at any cost; it’s no unnecessary lines.

How to get there

Start with a budget. For dependencies: we allow N new dependencies per quarter, and each needs a justification. For features: we don’t add without removing something or without a clear success metric. For code: we delete as much as we add. Budgets force tradeoffs and make “no” easier.

Then, make removal a first-class action. Sunset features that aren’t used. Remove dependencies that are redundant. Prune options that nobody touches. Removal is not failure; it’s maintenance. Schedule it.

Finally, default to “no.” When someone proposes an addition, the burden of proof is on them. What problem does it solve? What’s the cost? Can we solve it with what we have? Often the answer is yes, we just hadn’t looked.

When minimalism isn’t enough

There are domains where you need more: compliance, integration with legacy systems, or a broad feature set because the market expects it. Even then, minimalism is a direction. Within the required set, minimize. Don’t add “nice to have” on top of “must have” without a clear reason. And keep the core small: the kernel of the system should be as minimal as possible, with optional layers around it. That way you get the benefit of minimalism where it matters most.

Summary

Minimalism as engineering is the habit of counting the full cost of additions, dependencies, features, options, code, and adding only when the benefit is clear. It pays off in maintainability, debuggability, and team cognition. Get there with budgets, with removal as a first-class action, and with a default of “no.” Even when you can’t be fully minimal, minimize where it matters. Less is more when “less” means less to break, less to maintain, and less to think about.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleMinimalism as EngineeringRevAIssuedMar 5, 2025E-602

E-602record

sheet
E-602
title
Minimalism as Engineering
subtitle
Why less is more in system design: fewer dependencies, fewer features, fewer moving parts, and how to get there without sacrificing capability.
discipline
E · Essays
scale
1:1
revision
A
issued
Mar 5, 2025
refs
none
series
Essays
words
597

E-603Essays

The Architecture of Trust

How system design either builds or erodes trust, and what to optimize for when the user can’t see inside the black box.

Scale
1:1
Rev
A
Issued
Feb 20, 2025
Reading
3 min

Users don’t see your database or your API. They see behavior: Does it work? Is it fast? Does it fail in a way I understand? Trust is built or broken in those moments. This essay is about how architecture, the shape of your system, affects trust, and what to do about it.

Trust as experienced behavior

Trust is not a feature you add. It’s the result of repeated experience. If the system is reliable, fast, and predictable, trust grows. If it’s flaky, slow, or surprising, trust erodes. So the architecture of trust is the architecture of reliability, clarity, and consistency. Every component that the user depends on, the UI, the API, the data pipeline, contributes. A single point of failure, a silent error, or a confusing message can undo months of good behavior. Design for failure modes: what does the user see when something breaks? Do they get a clear message, a retry, or a generic error? Can they recover without calling support? The answers are architectural.

Transparency and explainability

When the system does something the user didn’t expect, a recommendation, a decision, a filter, trust depends on whether the user can understand why. That doesn’t mean every algorithm must be interpretable; it means the system should provide enough signal (e.g. “based on your past orders,” “because this item is out of stock”) that the user doesn’t feel in the dark. In AI systems, explainability is often framed as a technical problem (can we get saliency maps?). It’s also a product problem: what explanation does the user need to feel in control? Architecture should support that: logs, metadata, and UI hooks so that “why did this happen?” can be answered.

Consistency and promises

Trust is also about keeping promises. If you say “your data is private,” the architecture had better enforce that, encryption, access control, and no sneaky sharing. If you say “we’ll notify you when it’s done,” the system had better have a reliable notification path. Inconsistent behavior, sometimes it works, sometimes it doesn’t, is worse than consistently mediocre. So design for invariants: what do we promise, and how does the architecture guarantee it? Document those invariants and test for them.

Degradation and recovery

When things go wrong, trust is determined by how the system degrades and how it recovers. Graceful degradation (e.g. read-only mode when write is down, or a clear “we’re fixing this” message) preserves trust. Silent failure or data loss destroys it. Design for degradation paths: what’s the minimum useful behavior? What’s the recovery procedure? And design for communication: the user should know when something is wrong and when it’s fixed. Status pages, in-app messaging, and honest error copy are part of the architecture of trust.

Summary

The architecture of trust is the architecture of reliability, clarity, and consistency. Trust is built through repeated good behavior and eroded by flakiness, opacity, and broken promises. Design for failure modes and clear user-facing messages; support explainability where the user needs to understand why; enforce invariants that back up your promises; and design for graceful degradation and recovery. When the user can’t see inside the black box, the only thing they have is behavior. Make that behavior trustworthy.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleThe Architecture of TrustRevAIssuedFeb 20, 2025E-603

E-603record

sheet
E-603
title
The Architecture of Trust
subtitle
How system design either builds or erodes trust, and what to optimize for when the user can’t see inside the black box.
discipline
E · Essays
scale
1:1
revision
A
issued
Feb 20, 2025
refs
none
series
Essays
words
534

E-604Essays

Why Performance Is a Feature

Performance is not a nice-to-have. It’s a product requirement that affects conversion, retention, and trust. Here’s the case and the practice.

Scale
1:1
Rev
A
Issued
Feb 5, 2025
Reading
3 min

Performance is often relegated to “we’ll optimize later.” Later never comes, or it comes as a fire drill when users complain or when a competitor is faster. This essay makes the case that performance is a feature, one that should be specified, measured, and shipped like any other, and outlines how to treat it that way.

The case for performance as feature

Users notice slow. They might not articulate “LCP” or “TTI,” but they feel lag, they feel jank, and they leave. The data is clear: every 100ms delay in load time can cost conversion; every extra second increases bounce. On mobile and on slow networks, the effect is larger. So performance is not a technical detail; it’s a direct input to business outcomes. That makes it a feature: something we promise to the user. “This site loads quickly and responds when you tap” is a promise. We should design for it, test for it, and ship it.

Performance also affects perception of quality. A fast, smooth experience feels more trustworthy and more polished. A slow, janky one feels broken even if the content is good. So performance is part of the product’s personality. Treat it that way.

Specify it

If performance is a feature, it needs a spec. That means: define the metrics (e.g. LCP, FID, CLS, TTI), set targets (e.g. LCP under 2.5s on 4G, CLS under 0.1), and make them part of the definition of done. Without a number, “fast” is vague. With a number, you can measure, track, and regress. Put the targets in the product or engineering doc. Make them visible in dashboards and in CI.

Measure and enforce

Measurement has to be continuous. Use Lighthouse in CI and fail or warn when scores drop below the bar. Use real-user metrics (e.g. CrUX) when you can, so you see what users actually experience. Track trends: are we getting faster or slower over time? If a release regresses performance, it’s a bug, not a tradeoff, unless the tradeoff is explicit and agreed.

Enforcement means: no ship if we’re below the bar, unless there’s an exception and a plan to fix. That sounds strict, but it’s the only way to prevent the slow creep of “we’ll fix it later.” Performance regressions are easy to introduce and painful to fix in bulk. Catching them at merge time is cheap.

Design for performance from the start

The best way to hit performance targets is to design for them from the start. That means: a performance budget (so much JS, so much CSS, so many images), critical path awareness (what blocks first paint?), and a default of “small and fast.” If you add a heavy library or a big image, the budget should complain. If you block the main thread for hundreds of milliseconds, the metrics should show it. Making performance visible in the development loop prevents the “big rewrite to make it fast” later.

Summary

Performance is a feature: it affects conversion, retention, and trust, and it should be specified, measured, and enforced like any other feature. Specify with metrics and targets; measure continuously in CI and in the field; enforce by failing the build or blocking the ship when we’re below the bar. Design for performance from the start with budgets and critical-path discipline. When we do that, “we’ll optimize later” becomes “we’re already fast.”

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleWhy Performance Is a FeatureRevAIssuedFeb 5, 2025E-604

E-604record

sheet
E-604
title
Why Performance Is a Feature
subtitle
Performance is not a nice-to-have. It’s a product requirement that affects conversion, retention, and trust. Here’s the case and the practice.
discipline
E · Essays
scale
1:1
revision
A
issued
Feb 5, 2025
refs
none
series
Essays
words
562

E-605Essays

The Cost of Convenience

Why every dependency and every abstraction has a price, and how to pay it consciously.

Scale
1:1
Rev
A
Issued
Jan 28, 2025
Reading
4 min

Convenience is seductive. A library that does it for you, a framework that hides the details, a service that “just works”, each saves time today. The cost shows up later: in upgrade cycles, in debugging, in the day you need to do something the abstraction doesn’t support. This essay is about recognizing that cost and choosing convenience only when it’s worth it.

What convenience is

Convenience is the delegation of complexity. When you use a framework, you don’t write the router or the state machine; when you use a SaaS, you don’t run the servers. The trade is: you get speed and consistency now, and you accept the constraints and the ongoing relationship with the provider. There’s nothing wrong with that, as long as you’re aware you’re making the trade.

The problem starts when we treat convenience as free. It’s not. Every dependency is a commitment: to its API, its release cycle, its license, and its existence. When the maintainer stops updating, or the company pivots, or the license changes, you’re left holding the bag. The cost of convenience is often hidden and delayed. We’re bad at valuing delayed costs, so we over-consume convenience.

Where it bites

In software, convenience bites in a few classic ways. Upgrade debt: The more you depend on, the more you have to upgrade. Each upgrade can break things. So you delay, and then you’re many versions behind and the migration is painful. Lock-in: The abstraction that made the first 80% easy can make the last 20% impossible. You need to do something the library doesn’t support, and you’re stuck. Debugging: When something goes wrong, you’re debugging through layers you don’t own. Stack traces point into the framework; the bug might be in your usage or in the framework itself. Supply chain: A dependency can pull in dozens of transitive dependencies. One of them gets compromised or goes rogue, and you’re in the news. Convenience aggregates risk.

In life, the same pattern shows up: subscriptions, auto-renewals, default settings. We optimize for the moment of sign-up and forget the recurring cost and the hassle of exit. The cost of convenience is often a recurring cost and an exit cost.

How to pay consciously

First, make the cost visible. For every dependency or subscription, ask: What do I pay per year? What happens if I need to leave? What happens if the provider changes or disappears? Write it down. Second, set a budget. Decide how much “convenience” you’re willing to buy, in money, in lock-in, in future upgrade work. When something new comes along, check it against the budget. Third, prefer reversible choices. Prefer a library you can replace over one that’s wired into everything. Prefer a subscription you can cancel without penalty. Reversibility is a form of optionality; it’s worth something.

Fourth, build the habit of “could I do without this?” For small conveniences, the answer is often yes. You might not need that extra dependency; you might not need that subscription. Asking the question doesn’t mean you always say no, it means you only say yes when the benefit clearly outweighs the cost.

When convenience is worth it

Convenience is worth it when: (1) the domain is well-understood and the abstraction is stable (e.g. payment processing, auth); (2) the cost of building and maintaining your own is clearly higher than the cost of the dependency; (3) you have an exit plan or the dependency is easy to swap. It’s not worth it when: you’re early in exploration and the abstraction might be wrong; the dependency is huge and you need a tiny part of it; or the provider’s incentives don’t align with yours. In those cases, prefer less convenience and more control.

Summary

Convenience has a cost: upgrade debt, lock-in, debugging friction, and supply-chain risk. We underweight that cost because it’s delayed and diffuse. To pay consciously: make the cost visible, set a budget, prefer reversible choices, and ask “could I do without this?” Use convenience where the domain is stable and the trade is clear; avoid it where the abstraction might be wrong or the relationship is one-sided. The goal isn’t to avoid convenience, it’s to choose it with eyes open.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleThe Cost of ConvenienceRevAIssuedJan 28, 2025E-605

E-605record

sheet
E-605
title
The Cost of Convenience
subtitle
Why every dependency and every abstraction has a price, and how to pay it consciously.
discipline
E · Essays
scale
1:1
revision
A
issued
Jan 28, 2025
refs
none
series
Essays
words
702

E-606Essays

What Systems Thinking Actually Is

A long-form look at systems thinking: feedback loops, leverage points, and why it’s a discipline, not a buzzword.

Scale
1:1
Rev
A
Issued
Jan 10, 2025
Reading
5 min

Systems thinking is invoked everywhere, in product, in strategy, in engineering, but rarely defined with enough precision to be useful. This essay is an attempt to say what it is, why it matters, and how to practice it without falling into vagueness or cargo-cult diagrams.

Beyond the buzzword

When people say “we need to think in systems,” they often mean one of two things: (1) we need to consider how parts of the organization or product interact, or (2) we need to draw boxes and arrows. The first is necessary; the second is only useful if the diagram encodes real structure, feedback loops, delays, and stocks, and informs action. So let’s start with structure.

A system, in this sense, is a set of elements that interact over time. The elements can be teams, features, codebases, or physical assets. The interactions can be flows of information, money, or control. What makes it “systems thinking” is the focus on feedback: when A affects B and B affects A, directly or indirectly, you have a loop. Reinforcing loops amplify; balancing loops dampen. Most interesting behavior comes from the interplay of these loops and from delays, the time between cause and effect.

If you stop there, you have a vocabulary. The discipline is in using it: naming the loops in your own context, identifying where delay or misperception leads to overshoot or oscillation, and choosing interventions at leverage points, places where a small change can shift the behavior of the whole system. Donella Meadows’ “Leverage Points” is the canonical reference: the most powerful leverage points are often paradigm, goals, and system structure, not parameters or buffers. In practice, that means questioning the rules of the game and the measures of success, not only tuning the knobs.

Why it matters for builders

Engineers and product people are constantly making local decisions: add a feature, refactor a module, hire for a role. Each decision has downstream effects. Without a map of the system, you optimize for the wrong thing, e.g. shipping faster while accruing tech debt that will slow you later, or growing users while support capacity collapses. Systems thinking does not tell you the right answer; it helps you see the loops and delays so you can anticipate second-order effects and avoid surprises.

It also helps with communication. “We’re slow because we have a reinforcing loop: more features → more bugs → more firefighting → less time for design → more quick fixes → more bugs” is a story that can align the team. The diagram is a shared model. When everyone can point at the same loop and the same leverage point, you have a basis for prioritization and for saying no.

How to practice it

Start with one system you care about: your product development loop, your support pipeline, or your deployment flow. List the main elements and the main flows (what influences what?). Then ask: Where are the feedback loops? Is there a reinforcing loop that could run away (e.g. growth → overload → churn → need for growth)? Is there a balancing loop that’s too slow (e.g. we fix quality when we feel pain, but we feel pain only after release)? Name the delays: how long between “we ship” and “we see the bug”? Between “we hire” and “they’re productive”?

Next, look for leverage. Often the highest leverage is: change the goal (what we optimize for), change the rule (what we allow or require), or add or break a feedback loop (e.g. bring quality signals earlier). Parameter tweaks, more tests, more people, are lower leverage and often get eaten by the existing structure. That doesn’t mean ignore parameters; it means don’t expect them to fix a structural problem.

Finally, make it a habit. After a failure or a surprise, do a short retrospective: what loop was at play? What did we miss? Over time you build a repertoire of patterns (e.g. “this looks like a drift to low quality” or “this is a capacity trap”) and you get faster at both diagnosis and design.

Limits and pitfalls

Systems thinking is not a replacement for domain knowledge or for execution. It’s a lens. You can also overdo it: not every problem needs a full causal loop diagram. Use it when the problem is complex, when there are multiple stakeholders and delayed effects, and when local optimization has failed. And be humble: your model is wrong in places. Treat it as a hypothesis, test it with data and conversation, and update.

Another pitfall is using it to blame “the system” and avoid agency. The point is to find leverage, including in your own behavior and in the small changes you can make. Systems thinking should empower action, not paralyze it.

Summary

Systems thinking is the discipline of seeing feedback loops, delays, and leverage points in the structures we work within. It matters for builders because it improves decisions and alignment. Practice it by mapping one system at a time, naming loops and delays, and looking for high-leverage interventions. Use it when complexity and delay matter; don’t let it become vagueness or excuse-making. Done well, it’s one of the most practical tools for thinking clearly about cause and effect.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleWhat Systems Thinking Actually IsRevAIssuedJan 10, 2025E-606

E-606record

sheet
E-606
title
What Systems Thinking Actually Is
subtitle
A long-form look at systems thinking: feedback loops, leverage points, and why it’s a discipline, not a buzzword.
discipline
E · Essays
scale
1:1
revision
A
issued
Jan 10, 2025
refs
none
series
Essays
words
870

B-501Field Notes

Reinforcement Learning After the Reward Model

How reasoning models are actually trained now: verifiable rewards instead of learned ones, GRPO and its descendants, RL compute that scales on a curve, and agents trained inside environments. With the engineering lessons for anyone shipping LLM systems.

Scale
1:1
Rev
A
Issued
Oct 6, 2026
Reading
6 min

For a few years the story of post-training was RLHF: collect human preferences, train a reward model on them, optimise the policy against it with PPO. That recipe still exists, but it is no longer where the capability is coming from. The reasoning models of the last eighteen months were made by something simpler and, in a way, more honest. This note is a field survey of that shift, and of what it means for people who build on these models rather than train them.

The verifier replaces the judge

Reinforcement learning with verifiable rewards, RLVR, keeps the RL objective and throws away the reward model. Where a task has a checkable answer, a maths problem, a unit test, a constraint on the output format, a deterministic function decides whether the sampled completion is correct, and the policy is rewarded one or zero. The model explores many reasoning paths and the correct ones are reinforced; nothing is fitted to human taste, and nothing can be gamed except the verifier itself.

The surprising part is how much this unlocks. A 2025 study (Wen et al., RLVR Implicitly Incentivizes Correct Reasoning in Base LLMs) argued that RLVR does not merely find answers the base model could already produce by sampling: measured with a metric that credits correct intermediate reasoning, CoT-Pass@K, it extends the reasoning boundary on both maths and code. A 2026 line of evidence (Morris et al.) pushes the other way, suggesting RLVR works less by adding knowledge than by activating and reorganising capability already latent in the base model. Both can be true; what they agree on is that a binary, honest reward on a hard task changes the model's behaviour in ways supervised fine-tuning on gold solutions does not.

Work on extending this beyond checkable domains, such as RLPR, replaces the verifier with the model's own probability of the reference answer, so that the same machinery reaches tasks without a test to run.

The algorithms got cheaper, then careful

PPO needs a value network the size of the policy. GRPO (Shao et al., 2024) removed it: sample a group of completions for the same prompt, use the group's mean reward as the baseline, keep PPO's clipping. Half the memory, and a baseline that matches the structure of the problem.

GRPO's weaknesses showed up at scale, and 2025 was spent fixing them. DAPO decoupled the clipping range and raised the upper bound (clip-higher) to stop entropy collapsing, and sampled prompts dynamically so that batches are not filled with groups where every completion got the same reward and the gradient is zero. Dr. GRPO corrected a length bias in the original normalisation. GSPO moved the importance ratio from the token to the sequence. The family is now a standard menu, and most labs run a variant of it.

Then came the question of whether any of it scales. The Art of Scaling Reinforcement Learning Compute for LLMs (Khatri, Madaan, Agarwal and colleagues, October 2025) spent more than 400,000 GPU-hours answering it. Their findings: RL training follows a sigmoid in compute that can be fitted on small runs and extrapolated to large ones, which they validated on a run of 100,000 GPU-hours; not all recipes reach the same asymptote; and the details people argue about (loss aggregation, normalisation, curriculum, the off-policy scheme) mostly change how fast you get there, not where you end up. They packaged the best combination as ScaleRL, an asynchronous recipe built on GRPO. The practical point is that RL post-training can now be budgeted and forecast like pre-training, which is why frontier labs have been raising RL compute by an order of magnitude between model generations.

From answers to actions

The newest front is agentic RL: the policy does not produce an answer, it produces actions, tool calls, code, browser steps, and is optimised over a multi-turn episode against an executable environment. The 2025 survey The Landscape of Agentic Reinforcement Learning for LLMs frames it as the move from single-loop tool use to long-horizon decision making with state tracking and policy following.

The bottleneck has turned out to be the environments, not the algorithms. A 2026 survey of agentic environment engineering catalogues the problem: environments emulated by another language model hallucinate, drift, and lose state; real environments are slow and expensive to run at the scale RL needs. The response has been a wave of synthesised environments (ToolVerse, "agent world models" that generate task settings on demand), failure-driven training that mines where agents break (SENTINEL), and hybrid rewards that combine step-level and task-level signals so that a long episode is not scored only at its end. The field has, in effect, rediscovered that the training ground is the product.

What a builder should take from this

I do not train these models. I deploy them, in regulated settings, against databases and documents and people who have to sign off. Four lessons travel.

  1. Verifiers are an asset class. The labs got further with a binary checker than with a learned judge. The same is true of a production system: a deterministic check on a model's output (the SQL parses, the extracted field matches the schema, the number ties to the source) is worth more than a confidence score, and it is the thing an auditor will accept. Build the verifier before the prompt.
  2. Group-relative thinking applies to evaluation. GRPO's baseline is the other samples for the same prompt. The cheapest evaluation harness I know follows the same shape: sample several completions, compare within the group, and look hard at the prompts where they disagree. Those are the cases the model has not settled.
  3. Environments beat prompts. If agentic RL is limited by the fidelity of its environments, then the equivalent in deployment is the fidelity of staging. A staging stack that does not run the same contract as production is the LLM-emulated environment of your own organisation: it hallucinates success. This is why I insist on local parity.
  4. The compute curve is a planning tool. Once a capability follows a predictable curve, you can price it. That is now true of RL post-training, which means the gap between an open base model and a lab's reasoning model is, increasingly, a budget line rather than a secret.

The reward model was a stand-in for the thing we could not check. The lesson of the last two years is that it pays to find the thing you can.

Sources

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleReinforcement Learning After the Reward ModelRevAIssuedOct 6, 2026B-501

B-501record

sheet
B-501
title
Reinforcement Learning After the Reward Model
subtitle
How reasoning models are actually trained now: verifiable rewards instead of learned ones, GRPO and its descendants, RL compute that scales on a curve, and agents trained inside environments. With the engineering lessons for anyone shipping LLM systems.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Oct 6, 2026
refs
none
series
Field Notes
words
1182

B-502Field Notes

World Models, As Issued: The State of the Set in October 2026

A field survey of where world models actually stand: V-JEPA 2.1, LeJEPA, Dreamer 4, Genie 3, Marble and Atlas, Cosmos, and a billion-dollar bet in Paris. What shipped, what it cost, and what it is for.

Scale
1:1
Rev
A
Issued
Oct 6, 2026
Reading
7 min

"World model" has become the phrase every lab reaches for, and it now covers at least three different things. This note is a survey of what has actually been issued, with dates and figures, sorted by what each system is for. I have kept to what the papers and announcements say and marked where I am inferring.

Three kinds of world model

It helps to separate them first.

  1. Latent predictors. The model learns a representation of the world and predicts how that representation evolves. It never draws anything. Its purpose is planning and control. This is the JEPA line.
  2. Generative simulators. The model renders what the world will look like, frame by frame, often conditioned on actions. Its purpose is to be a training ground for agents, or to be looked at. Dreamer, Genie and Cosmos live here.
  3. Spatial generators. The model produces an explorable 3D scene from a prompt or a few images. Its purpose is content, simulation and, increasingly, robot training data. World Labs lives here.

The three overlap at the edges, and the most interesting 2026 systems sit on an edge.

Latent predictors

V-JEPA 2 (Meta, June 2025) was the first video-trained world model to show zero-shot robot planning. It is a 1.2-billion-parameter model pretrained on more than a million hours of internet video and about a million images; an action-conditioned variant, V-JEPA 2-AC, was then trained on roughly 62 hours of robot data and used for model-predictive control on a real arm, in a new lab, with no task-specific training.

V-JEPA 2.1 (March 2026) is the same family with a denser objective. Every token is supervised, visible and masked alike; self-supervision is applied at intermediate encoder layers, not only the last; images and video get separate tokenizers feeding one shared encoder. The paper reports a 20-point improvement in real-robot grasping success over V-JEPA 2-AC, and state-of-the-art figures on short-term object interaction, action anticipation, depth estimation and a robotic navigation benchmark.

LeJEPA (Balestriero and LeCun, November 2025) is a training method, not a model. It identifies the isotropic Gaussian as the embedding distribution that minimises downstream risk, and enforces it with a regulariser, SIGReg, that works through random one-dimensional projections in linear time. The result is JEPA training with no teacher-student pair, no stop-gradient, one hyperparameter, and a loss that correlates with linear-probe accuracy closely enough to select models without labels. They report stable runs up to a 1.8-billion-parameter ViT. Follow-on work in 2026 has already proposed variants of the regulariser and applied the recipe to EEG.

AMI Labs (Paris, launched March 2026) is where this line is now being pursued commercially. Yann LeCun, who left Meta, is chairman; Alexandre LeBrun is CEO. The seed round was $1.03 billion at a $3.5 billion pre-money valuation, led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital and Bezos Expeditions, with Nvidia and Samsung among the other backers. The stated plan is JEPA-based world models that learn from sensory data rather than text, and a commitment to open-sourcing much of the code. LeBrun has said commercial applications are some way off.

Generative simulators

Dreamer 4 (Hafner, Yan and Lillicrap, 2025) is the clearest result of the year. The agent is trained entirely inside a learned world model of Minecraft, with no interaction with the real game, and is the first to reach diamonds from offline data alone, a task that needs more than twenty thousand mouse and keyboard actions from raw pixels. It outperforms OpenAI's VPT offline agent with a hundred times less action-labelled data, and the world model runs in real time on a single GPU. The world model is trained on a large corpus of unlabelled gameplay video with actions supplied for only a small subset; the project page frames the insight as most of the world knowledge coming from the unlabelled footage.

DreamerV3, its predecessor, was published in Nature in April 2025: one configuration across more than 150 tasks, and the first agent to collect Minecraft diamonds from scratch with online interaction.

Genie 3 (Google DeepMind, August 2025) generates interactive worlds from a text prompt at 720p and 24 frames per second, with roughly a minute of memory, and responds to the user's movement in real time. In January 2026 DeepMind opened it to AI Ultra subscribers as Project Genie. It is a simulator built to be inhabited; the obvious next use is as an environment for training agents, and the research line is heading there.

NVIDIA Cosmos is a platform of world foundation models and data tools aimed at physical AI: synthetic, physics-aware video for training robots and autonomous vehicles. NVIDIA has reported two million downloads. Its purpose is upstream of the others: it exists to manufacture training data.

Spatial generators

World Labs shipped Marble in late 2025, a model that builds explorable 3D scenes from text, images or video, and followed it with Marble 1.1 and 1.1 Plus, which improved lighting and artefacts and allowed larger scenes. On 1 September 2026 it announced Atlas, described as a multimodal world model that generates image and video frames with exact camera control and reconstructs them in 3D, released to selected partners by application.

Then, on 28 September 2026, AMD agreed to acquire World Labs in an all-stock transaction valued at about $8.2 billion, expected to close by the end of the year subject to approvals. Fei-Fei Li joins AMD as executive vice president and chief scientist, reporting to Lisa Su; the stated rationale is to tie model research to hardware and systems design. It is the largest transaction in the field so far, and it says something about where chip companies think the next workload is.

What it adds up to

A few readings, mine rather than the papers'.

  • The split is purpose, not technique. Latent predictors are for choosing actions; generative simulators are for training and for looking at; spatial generators are for content and data. Arguments about which is "the" world model miss that they are different tools.
  • Unlabelled footage is the asset. Dreamer 4 and V-JEPA 2 both draw most of their knowledge from video with no actions attached, and bolt on a small labelled set at the end. The economics of robot learning follow from this.
  • The money has moved. A billion-dollar seed in Paris and an $8.2 billion acquisition in Santa Clara in the same year, for two companies that have shipped research models and an early product. The bet is on the category.
  • Nothing here is a product you can buy for your own problem yet, with the partial exception of Cosmos as a data tool. If you are building a system today, world models are a reason to keep your pipelines modular, not a component to order.

Sources

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleWorld Models, As Issued: The State of the Set in October 2026RevAIssuedOct 6, 2026B-502

B-502record

sheet
B-502
title
World Models, As Issued: The State of the Set in October 2026
subtitle
A field survey of where world models actually stand: V-JEPA 2.1, LeJEPA, Dreamer 4, Genie 3, Marble and Atlas, Cosmos, and a billion-dollar bet in Paris. What shipped, what it cost, and what it is for.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Oct 6, 2026
refs
none
series
Field Notes
words
1217

B-503Field Notes

Experimental UI as an Engineering Discipline

Why treating experimental UI and creative front-end as engineering, not “designer code”, raises the bar for both art and performance.

Scale
1:1
Rev
A
Issued
Mar 8, 2025
Reading
2 min

Experimental UI, custom layouts, generative art, spatial navigation, is often dismissed as eye candy or “designer code.” When it’s done as engineering, it becomes repeatable, maintainable, and fast. This post is about how to do it that way.

Constraints as creativity

The best experimental UI works within constraints: a performance budget, accessibility requirements, and a clear contract with the rest of the app. Constraints force choices. “We have 50 KB for this interaction” leads to smarter techniques than “we’ll optimize later.” Treat the budget and a11y as part of the brief, not an afterthought.

Technique over trend

Trends (e.g. glassmorphism, parallax everywhere) age badly. Technique, how you use transform and opacity, how you structure your animation loop, how you lazy-load and dispose, lasts. Invest in understanding the platform: compositor, requestAnimationFrame, IntersectionObserver, and the visibility API. That knowledge transfers to every project.

Measurable outcomes

Experimental UI should still have success criteria: Lighthouse score, CLS, time to interactive, and keyboard/screen-reader usability. If the experience is beautiful but fails accessibility or performance, it’s incomplete. Define “done” up front and measure.

Reusability and abstraction

Even one-off experiences benefit from abstraction: a small motion library, a shared pattern for “reveal on scroll,” or a consistent way to pause animations when the tab is hidden. That reduces bugs and makes the next experiment faster. Experimental UI as engineering means building a toolkit, not only a single page.

Summary

Experimental UI as engineering means: constraints as part of the brief, technique over trend, measurable outcomes (performance and a11y), and reusable patterns. When you treat it that way, the result is both expressive and robust.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleExperimental UI as an Engineering DisciplineRevAIssuedMar 8, 2025B-503

B-503record

sheet
B-503
title
Experimental UI as an Engineering Discipline
subtitle
Why treating experimental UI and creative front-end as engineering, not “designer code”, raises the bar for both art and performance.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Mar 8, 2025
refs
none
series
Field Notes
words
272

B-504Field Notes

Architecture Design for Healthcare AI

How to design AI systems in healthcare: compliance, safety, and the primacy of the clinician in the loop.

Scale
1:1
Rev
A
Issued
Mar 1, 2025
Reading
2 min

Healthcare AI is not “AI plus healthcare.” It is a domain where errors can harm patients, data is highly regulated, and the end user is often a clinician who must trust and override the system. Architecture has to reflect that from day one, whether you are extracting structured variables from EHRs at thousands of documents per day (clinical NLP), running multimodal pipelines under HIPAA-aware AWS (finance & health agents), or grounding programme KPIs so LLM narratives never invent counts (AIDA).

Safety and accountability

The system must support accountability. That means: every recommendation or prediction should be traceable to the model version and the data slice it used. Audit logs, versioned models, and the ability to explain (at least at a high level) why the system said what it said are not nice-to-haves. They are part of the product. Design for auditability in the data pipeline and in the inference path.

Clinician in the loop

Automation that cannot be overridden or ignored is dangerous. The clinician must be the final decision-maker. The UI and the API should make it easy to see the model’s output, the confidence (if you have it), and to dismiss or correct. Design for “assist, don’t replace” and for graceful degradation when the model is uncertain or the input is out of distribution.

Compliance and data

HIPAA, GDPR, and local regulations constrain where data lives, how it’s transmitted, and how long it’s retained. Architecture choices, on-prem vs. cloud, which cloud, how data is tokenized or de-identified, depend on the jurisdiction and the use case. Build with compliance in mind: encrypt in transit and at rest, minimize retention, and document data flows. Assume you will be audited.

Reliability and fallbacks

Healthcare workflows often run 24/7. The AI component should fail gracefully: if the model is down or slow, the system should fall back to a safe default (e.g. no recommendation, or a generic one) and surface the failure. No silent wrong answers. Design for observability so that outages and errors are detected and escalated.

Summary

Healthcare AI architecture prioritizes safety, accountability, clinician-in-the-loop, compliance, and reliability. The model is one component; the surrounding system, data pipeline, API, UI, and ops, must be designed for the domain. Get that right before chasing accuracy alone.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleArchitecture Design for Healthcare AIRevAIssuedMar 1, 2025B-504

B-504record

sheet
B-504
title
Architecture Design for Healthcare AI
subtitle
How to design AI systems in healthcare: compliance, safety, and the primacy of the clinician in the loop.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Mar 1, 2025
refs
none
series
Field Notes
words
379

B-505Field Notes

AI in Restaurant Automation

Where AI actually helps in restaurant operations, order prediction, inventory, and kitchen flow, and where it’s hype.

Scale
1:1
Rev
A
Issued
Feb 25, 2025
Reading
2 min

While founding engineering on a restaurant SaaS pilot (200+ sites), the wins came from the back of the house, not chatbots at the register. Restaurants run on thin margins and chaotic inputs: weather, events, no-shows, and daily specials. AI can help in narrow, well-defined areas. It can also waste time and money if applied where the problem is process, not prediction. This post is about the former; the restaurant forecasting case study covers what we shipped.

Where AI helps

Demand and prep: Predicting covers or order mix for the next service lets the kitchen prep the right amount. That’s a classic forecasting problem: historical data, maybe some exogenous signals (day of week, events), and a model that’s updated regularly. The value is less waste and fewer stockouts. The key is defining the horizon (same day, next day) and the granularity (by dish or by category).

Inventory and ordering: Similar idea: predict what you’ll need and suggest order quantities. The constraint is lead time and minimum order sizes. The model has to respect those and surface confidence so humans can override.

Kitchen flow: Optimizing sequence or timing of tickets is harder, the state space is large and the cost of being wrong (late food, wrong order) is high. Here AI is more assistive: suggest a sequence, but let the expediter decide. Full autonomy in the kitchen is a long way off.

Where it’s hype

Replacing the human at the register or in the dining room with a chatbot is often a solution in search of a problem. So is “AI-powered menus” that don’t change the underlying economics. The real gains are in the back: forecasting, inventory, and decision support. Front-of-house automation that annoys customers or adds friction is a net negative.

Implementation notes

Data quality is the bottleneck. You need consistent point-of-sale and inventory data, and you need to align it with outcomes (waste, stockouts, satisfaction). Start with one workflow, e.g. prep list for the next day, and prove value before expanding. Use simple models first; complexity is rarely justified in this domain. And always have a fallback: when the model is down or wrong, the restaurant should still run.

Summary

AI in restaurant automation is most valuable in the back: demand forecasting, inventory, and prep. Front-of-house and “AI everywhere” are often hype. Focus on one workflow, good data, and a human in the loop.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAI in Restaurant AutomationRevAIssuedFeb 25, 2025B-505

B-505record

sheet
B-505
title
AI in Restaurant Automation
subtitle
Where AI actually helps in restaurant operations, order prediction, inventory, and kitchen flow, and where it’s hype.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Feb 25, 2025
refs
none
series
Field Notes
words
398

B-506Field Notes

Systems Thinking for Founders

Why founders who think in systems ship better products and make fewer irreversible mistakes.

Scale
1:1
Rev
A
Issued
Feb 18, 2025
Reading
2 min

Founders are rewarded for speed. But speed without structure leads to tech debt, unclear ownership, and decisions that are hard to undo. Systems thinking is the habit of seeing your product and company as a set of interacting parts with feedback loops, bottlenecks, and leverage points. It makes you faster in the long run.

What is systems thinking?

Systems thinking is the discipline of mapping cause and effect across boundaries: engineering, product, go-to-market, support. It asks: What are the key loops? Where does information get stuck? What happens if we change this variable? You don’t need formal modeling; you need the reflex to draw the diagram and name the feedback (reinforcing or balancing) and the delay.

Why it matters for founders

Founders set the initial conditions. A messy repo, no runbooks, or “we’ll document later” compounds. So does clear ownership, a single source of truth for metrics, and a culture of “we can explain why we built it this way.” Systems thinking helps you see which early decisions will compound and which will fade. It also helps you communicate with investors and hires: “Here’s how the pieces fit” is more convincing than “we’re doing a bit of everything.”

Practical habits

  • Map the value flow: From user action to backend to data to decision. Where does value get created and where does it leak?
  • Name the bottlenecks: Is it engineering capacity, distribution, or something else? Don’t optimize the wrong lever.
  • Design for reversibility: Can you roll back a feature? Can you sunset a dependency? If not, you’ve made an irreversible bet; treat it explicitly.
  • Feedback loops: How do you know if a change worked? Define the metric and the cadence before you ship.

Summary

Systems thinking for founders is not academic. It is the difference between “we moved fast” and “we moved fast in the right direction.” Invest in the habit early; the payoff is fewer surprises and more durable speed.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleSystems Thinking for FoundersRevAIssuedFeb 18, 2025B-506

B-506record

sheet
B-506
title
Systems Thinking for Founders
subtitle
Why founders who think in systems ship better products and make fewer irreversible mistakes.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Feb 18, 2025
refs
none
series
Field Notes
words
328

B-507Field Notes

Performance Engineering on Static Hosting

How to hit Lighthouse 95+ and sub-second FCP when you have no server: budgets, critical path, and discipline.

Scale
1:1
Rev
A
Issued
Feb 10, 2025
Reading
2 min

Static hosting, GitHub Pages, Netlify, Cloudflare Pages, gives you global edge delivery and no server to maintain. It also means every byte and every request is visible to the user. Performance is not optional; it is the product. Here is how to treat it as engineering.

Budgets first

Set a performance budget before you build. For example: JS under 70 KB (gzipped), CSS under 20 KB, images optimized and lazy-loaded. Make the build fail or warn when the budget is exceeded. That forces tradeoffs up front instead of “we’ll fix it later.”

Critical path

The only thing that must block first paint is the HTML and the minimal CSS needed above the fold. Fonts should load with font-display: swap or optional; non-critical CSS can be deferred or inlined only for the first screen. Scripts should be deferred or loaded as modules so parsing does not block rendering. Measure FCP and LCP on real devices and on 3G; if it’s slow, something is on the critical path that does not need to be.

Static does not mean dumb

Use preloading for key resources (e.g. LCP image, main font). Use loading="lazy" for below-the-fold images. Preconnect to origins you will hit (e.g. fonts, API). All of this is static configuration; no server required. Sitemaps and meta tags are free; use them for SEO and crawlers.

No layout thrashing

If you add client-side interactivity, avoid layout thrashing: batch reads and writes, use requestAnimationFrame for visual updates, and prefer transform and opacity for animations so the compositor does the work. Respect prefers-reduced-motion and pause or simplify animations when the user prefers it.

Measure and enforce

Run Lighthouse in CI. Set a minimum performance score (e.g. 95) and fail the build if it drops. Use real-user metrics if you can (e.g. CrUX) to catch regressions in the field. Performance engineering on static hosting is mostly discipline: budgets, critical path, and measurement. Do it from day one.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitlePerformance Engineering on Static HostingRevAIssuedFeb 10, 2025B-507

B-507record

sheet
B-507
title
Performance Engineering on Static Hosting
subtitle
How to hit Lighthouse 95+ and sub-second FCP when you have no server: budgets, critical path, and discipline.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Feb 10, 2025
refs
none
series
Field Notes
words
325

B-508Field Notes

Building SaaS with a Zero-Dependency Mindset

Why minimizing dependencies is a strategic advantage in SaaS, and how to do it without reinventing the wheel.

Scale
1:1
Rev
A
Issued
Feb 1, 2025
Reading
2 min

“Zero dependency” does not mean you write every line yourself. It means you treat every dependency as a cost and only pay when the benefit is clear. For SaaS, that mindset reduces supply-chain risk, keeps bundles small, and makes upgrades and audits tractable.

Dependency as liability

Every dependency can break, change license, or go unmaintained. In Node, left-pad showed how a tiny package can take down the ecosystem. In front-end, heavy frameworks and UI libraries lock you into their release cycle and their bugs. The more you depend on, the more you have to track, test, and upgrade. So the default should be: do we need this?

When to add a dependency

Add a dependency when: (1) the problem is well-defined and the library solves it correctly, (2) the maintenance burden of doing it yourself is higher than the burden of upgrading the library, and (3) the license and security posture are acceptable. For example: cryptography, date/time edge cases, and complex parsing are often worth a dependency. A button component or a shallow clone utility usually is not.

How to minimize

  • Prefer the platform: use the Fetch API, CSS Grid, native modules. You get updates with the runtime.
  • Prefer small, focused packages over frameworks when you can. One function to do one thing is easier to replace than a whole ecosystem.
  • Audit regularly: run npm outdated, check for security advisories, and remove what you no longer use.
  • Abstract the dependency behind your own interface. If you ever need to swap or remove it, you have one place to change.

Zero-dependency as strategy

In practice, “zero dependency” is a direction, not a literal count. The goal is to own your critical path and to keep optional dependencies optional. For a content site, that might mean static HTML and a tiny script. For a SaaS app, it might mean a minimal runtime and a small set of well-chosen libraries. The mindset is: every dependency should earn its place. That keeps the system understandable, fast, and under your control.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleBuilding SaaS with a Zero-Dependency MindsetRevAIssuedFeb 1, 2025B-508

B-508record

sheet
B-508
title
Building SaaS with a Zero-Dependency Mindset
subtitle
Why minimizing dependencies is a strategic advantage in SaaS, and how to do it without reinventing the wheel.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Feb 1, 2025
refs
none
series
Field Notes
words
344

B-509Field Notes

DNS-Level Ad Blocking. System Design

How to design a DNS-based ad blocking system: architecture, tradeoffs, and what you give up when you block at the resolver.

Scale
1:1
Rev
A
Issued
Jan 22, 2025
Reading
3 min

Blocking ads and trackers at the DNS layer is elegant: one place to enforce policy, no per-app logic, and visibility into every name your devices resolve. But the system design is non-trivial. This post outlines how to think about it.

Why DNS?

DNS is the first step in almost every connection. If you block or redirect a hostname at the resolver, the request never reaches the ad server. That gives you a single point of control, works across devices and apps, and avoids the complexity of per-browser or per-OS extensions. The tradeoff is granularity: you operate at the level of hostnames, not URLs or page elements. For many ads and trackers, that is enough.

Architecture components

A minimal system has: a resolver (your DNS server), a blocklist (domains to block or sinkhole), and a way to get DNS queries to your resolver (router, VPN, or device config). Optional but valuable: logging (for debugging and analytics), allowlisting (to unbreak sites), and encrypted transport (DoH/DoT) so queries are not visible on the wire.

The resolver can be something like Pi-hole, AdGuard Home, or a custom stub that forwards to an upstream and filters responses. The blocklist is the product of community efforts (e.g. OISD, Steven Black’s list) plus your own rules. Keep the list in a format that your resolver understands and version it like code.

Tradeoffs

Latency: Every query can be checked against the list. Use efficient data structures (trie or hash set) and keep the list in memory. A few milliseconds per query is acceptable for most users.

Accuracy: Blocklists can overblock (breaking sites) or underblock (missing trackers). Allowlisting and regular list updates are essential. Prefer lists that are maintained and documented.

Privacy: If you log queries, you have sensitive data. Decide retention and access. Prefer not logging by default; add logging only where needed for debugging.

Single point of failure: If your resolver is down, clients may fall back to the ISP resolver or fail. Design for fail-open or fail-closed based on your risk model. Many home setups fail-open so the internet still works when the Pi-hole is off.

Metrics that matter

  • Block rate: share of queries that hit a blocked domain.
  • False positives: sites or services that break and need allowlisting.
  • Resolver latency (p50, p99).
  • List size and update frequency.

Summary

DNS-level ad blocking is a single point of control with clear tradeoffs: hostname-level granularity, need for allowlisting, and operational responsibility for the resolver. Design for latency, accuracy, and privacy from the start, and treat the blocklist as a maintained artifact.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleDNS-Level Ad Blocking. System DesignRevAIssuedJan 22, 2025B-509

B-509record

sheet
B-509
title
DNS-Level Ad Blocking. System Design
subtitle
How to design a DNS-based ad blocking system: architecture, tradeoffs, and what you give up when you block at the resolver.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Jan 22, 2025
refs
none
series
Field Notes
words
432

B-510Field Notes

AI Infrastructure Philosophy

Why the best AI systems treat infrastructure as a first-class product, and how to think about reliability, scale, and operator experience.

Scale
1:1
Rev
A
Issued
Jan 15, 2025
Reading
3 min

Infrastructure is not the boring part of AI. It is the part that determines whether your model ever runs in production, whether it degrades gracefully, and whether your team can iterate without setting the building on fire. Forward-deployed work (shared LLM gateways on AWS CDK, Kafka paths under 200ms, governed NL2SQL with human confirmation) is where that philosophy stops being abstract; see the multi-tenant fertility platform and video intelligence case studies for concrete tradeoffs.

Infrastructure as product

Most teams treat infrastructure as a cost center: something to minimize, outsource, or ignore until it breaks. The better framing is product. Your inference pipeline, your feature store, your evaluation loop, these are products with users (other services, data scientists, and ultimately the end user). They need contracts, versioning, and clear failure modes.

When you design AI infrastructure as product, you start asking: What is the API? What are the SLAs? What happens when the model is stale? What happens when the GPU is OOM? The answers become part of the design, not an afterthought.

Reliability over novelty

It is tempting to chase the latest model or the cleverest fine-tuning trick. In production, what matters more is: Does it run every time? Can we roll back? Can we A/B test without redeploying the world? Reliability is a feature. Latency budgets, retry policies, and fallback paths are not bureaucracy, they are the difference between a demo and a system people depend on.

Scale is a function of design

Scale is not “add more machines.” It is “design so that adding more machines (or shrinking to one) is a configuration change.” That means stateless services, idempotent jobs, and data flows that can be partitioned. It also means accepting that some workloads are fundamentally single-node until you change the algorithm. Honest scaling is about matching the architecture to the access pattern, not the other way around.

Operator experience

The best AI systems are operable: they expose enough telemetry to debug, enough knobs to tune, and enough documentation to onboard. If only one person can run the pipeline, it is not infrastructure, it is a script. Invest in runbooks, structured logs, and clear ownership. The team that builds the model should be able to support it in production, and that only works if the infrastructure is legible.

Summary

AI infrastructure philosophy, in short: treat it as product, prioritize reliability, design for scale from first principles, and make it operable. Everything else is implementation detail.

Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleAI Infrastructure PhilosophyRevAIssuedJan 15, 2025B-510

B-510record

sheet
B-510
title
AI Infrastructure Philosophy
subtitle
Why the best AI systems treat infrastructure as a first-class product, and how to think about reliability, scale, and operator experience.
discipline
B · Field Notes
scale
1:1
revision
A
issued
Jan 15, 2025
refs
none
series
Field Notes
words
412

C-700Correspondence

Correspondence

Open a line

Let's build something that matters.

Scale
NTS
Rev
B
Issued
2026.09
1Transmittal, in elevationC-700 · NTS
Email
hello@akshaybajpai.com
LinkedIn
linkedin.com/in/ax5hay
GitHub
github.com/ax5hay
Write
www.akshaybajpai.com/contact/
Architecture of IntelligenceDrawing set · Akshay Bajpai · akshaybajpai.comTitleCorrespondenceRevBIssued2026.09C-700

C-700record

sheet
C-700
title
Correspondence
subtitle
Open a line
discipline
C · Correspondence
scale
NTS
revision
B
issued
2026.09
refs
none