AI and ML that ship to production — not demos.
This is where most of the demand is today, and it’s where Leopard Data spends most of its time: LLM-powered products and AI agents built on Claude, AI-driven cloud migration that ports real codebases, and distributed machine learning — deep-learning models and feature-ranking engines — engineered to run at enterprise scale on Kubernetes.
What We Do in AI & ML
Four areas, all backed by shipped, production work.
AI-First Product Engineering
LLM features built on Anthropic Claude — multi-model routing (Haiku / Sonnet / Opus), RAG pipelines, MCP servers and structured tool use, prompt caching, and token-aware server-side billing so AI features pay for themselves.
AI Agents & Agentic Migration
Agents that do real work — including agentic code migration that analyzes and ports large codebases between clouds, running Claude on AWS Bedrock with custom MCP servers and a human architect reviewing every pass.
Distributed ML & Deep Learning
LSTM and DeepAR in PyTorch, feature-ranking engines over 200K+ features, and the hard part most teams skip: distributing billions of calculations across Kubernetes with Ray and KEDA so the models actually finish.
ML in Production .NET
ML.NET-powered forecasting and scoring inside enterprise .NET applications — deterministic outputs from probabilistic inputs, with reproducible training pipelines, unit tests, and drift monitoring.
The Proof
Real engagements and a shipped, AI-first SaaS platform.
AI Agents Migrating a Healthcare Platform from AWS to GCP
An AI migration agent — Claude on Bedrock plus custom MCP servers — that analyzes and ports large FHIR/HL7 TypeScript codebases across 350+ code bases to Google Cloud.
Read the case study Industrial · Distributed MLDistributed Machine Learning at Scale
A feature-ranking engine over 200K+ features and an LSTM/DeepAR prediction engine, parallelized across Kubernetes with KEDA and Ray.
Read the case study Industrial · AI & MLKoch Industries — AI Production Forecasting
Distributed AI / ML platform with REST services and parallelized feature processing on AWS EKS, forecasting capacity across multiple business units.
Read the case study How We Build AI Into Products · ExampleInside GMI — Claude & ML.NET in Production
A worked example of how we integrate AI into apps and products, with architecture diagrams: deterministic ML.NET grading, a Claude file-analysis layer, and token-aware cost caps — the real services and numbers.
See how we integrate AI into productsThe AI & ML Stack
The tools behind the work.
LLMs & Agents
Machine Learning
Distributed Compute
AI Governance for Regulated Industries
Built for healthcare and financial-services constraints — where AI has to be safe, auditable, and cost-controlled.
Your data stays yours
Claude and other models run in-tenant through AWS Bedrock, Azure OpenAI, and Google Vertex AI. Your prompts and documents are never used to train foundation models, and inference stays inside your own cloud account and VPC.
PHI & PII discipline
Data minimization, redaction, and scoping so regulated data never leaves controlled boundaries — the same discipline applied to real FHIR/HL7 healthcare interoperability work on multi-tenant platforms.
Auditability & cost control
Every model call is metered server-side with token accounting, logging, and hard spend caps — the same controls running in our own SaaS, so AI spend never surprises finance.
Human in the loop
An architect reviews every agent pass. Structured outputs, eval gates, and deterministic ML.NET for anything that must be repeatable catch hallucinations before they reach production.
Working with AI — Straight Answers
The questions enterprise teams ask before putting AI into production.
Will our data be used to train the model?
No. On AWS Bedrock, Azure OpenAI, and Google Vertex AI, your prompts and documents stay in your cloud tenant and are not used to train the underlying models. We design pipelines so sensitive data never leaves your controlled environment.
How do you keep AI costs predictable?
Model tiering (Haiku / Sonnet / Opus per task), prompt caching, retry/backoff, request staggering, and server-side token metering with hard spend caps. Users see an estimated cost before a job runs — exactly how billing works in Grade My Investments.
How do you handle hallucinations and reliability?
Structured, tool-constrained outputs; eval pipelines that prove a prompt before it ships; deterministic ML.NET for anything that must be repeatable; and a human architect reviewing agent output. AI accelerates the work — it doesn’t get the final word unchecked.
When should we not use AI?
When a deterministic rule, a SQL query, or classic ML is cheaper and more reliable. We tell you where AI earns its keep and where it just adds cost and risk — that judgment is the point of hiring a senior architect instead of a model.
Claude, GPT, Gemini, or open-weight — which model?
Vendor-neutral. We pick per task on cost, latency, and reasoning depth. Most of our production work runs on Anthropic Claude, but we integrate GPT, Gemini, Mistral, or open-weight models wherever they fit better.
Can you work inside our cloud and compliance boundaries?
Yes. We deploy against your Azure, AWS, or Google Cloud accounts — inside your VPC and IAM, managed as infrastructure-as-code with audit logging — rather than shipping your data to a third-party tool.
Putting AI or ML into production?
From LLM features and AI agents to distributed deep learning, Leopard Data ships the real thing. Corp-to-Corp engagements out of Plano, TX.