Independent AI engineer · Available for consulting

Ship AI that survives production.

I'm Suyash Dubey — an applied, forward-deployed AI engineer. Three years building agentic systems and retrieval pipelines in healthcare and public digital infrastructure, where a wrong output is a clinical risk rather than a bad demo.

LangGraphFastAPIAWS BedrockLangfuseKubernetes

Grounded retrieval on Every call traced on Eval gate before ship PII / PHI redaction enforced
0%

less clinician documentation time after the scribe agent rolled out

0%

accuracy lift on real-time EEG signal classification

0%

manual steps removed from a content operations pipeline

0 yrs

shipping LLM systems in regulated, data-sensitive domains

Figures are approximate and reflect internal measurement at the time, in previous roles. Happy to walk through how each was measured on a call.

What I do

Services built
for shipping

Four ways to work together. Every one of them ends in something running, instrumented and handed over — not a document about what could be built.

01

AI MVP development

Zero to one. We agree the single outcome the product has to prove, then I build and deploy the thinnest system that proves it — API, model orchestration, retrieval, front end if you need one, and the evaluation that tells you whether it actually works.

0 → 1Deployed, not demoedFixed scope
Typical 4–6 weeksEnds in a live system
02

Agentic systems & RAG

Stateful LangGraph workflows with conditional branching, tool use and human-in-the-loop review, so the agent handles the messy middle instead of escalating to a person on every turn. Retrieval that is grounded in your own corpus, with the citation trail to prove it.

LangGraphGrounded RAGTool useHITL
Typical 3–8 weeksEnds in a traced pipeline
03

AI strategy & forward-deployed consulting

I sit with the people who will actually use the system, in their language, and come back with what is feasible, what it costs, what it risks and the order to build it in. Includes translating AI-governance and security requirements into concrete technical controls rather than policy prose.

Discovery sprintArchitectureCosted roadmap
Typical 1–2 weeksEnds in a decision you can fund
04

Evaluation, observability & governance

The part most teams bolt on too late. Evaluation harnesses that gate releases, Langfuse tracing on every model call, versioned prompts, PII and PHI redaction before inference, and dashboards for token spend, latency and quality so regressions surface before your users find them.

Eval harnessesLangfusePrompt governancePHI redaction
Typical 2–4 weeksEnds in release gates

How it runs

The forward-
deployed loop

The same four steps whether it is a two-week discovery or a six-month build. The point is to get real usage in front of the system as early as it can survive it.

01

Scope in plain language

I talk to the people who will use it, not only the people who commissioned it, and write the behaviour spec in their words.

02

Build the thin slice

One end-to-end path through the system, deployed. Enough to be wrong in public and cheap to change.

03

Instrument before scaling

Traces, evals, redaction and cost telemetry go in on the first commit — never as a pre-launch scramble.

04

Iterate on real usage

Weekly cycles against what people actually did with it, with the evaluation suite deciding what ships.

Selected work

Systems that
went live

Other builds

Compliance audit agent

Agentic API that ingests a URL, scrapes it autonomously, judges it against a written policy with an LLM, and returns structured non-compliant findings with justification.

LangChain · Flask · Pydantic

RAG document intelligence API

End-to-end retrieval-augmented QA over unstructured documents — embedding retrieval, grounded generation and an evaluation harness, exposed as REST.

LangChain · FAISS · FastAPI

WhatsApp virtual try-on

Zero-human-in-the-loop automation: inbound WhatsApp message via Twilio, diffusion try-on model, result delivered straight back to the sender.

Flask · Twilio · Gradio

In-memory-compute NN accelerator

Evaluated LeNet, AlexNet, VGG and GoogLeNet for in-memory computation substitution, cutting average execution time by around three minutes.

TensorFlow · Keras · Simulink

Track record

Where I've
shipped

Healthcare, public digital infrastructure, research and content platforms — consistently the environments where a wrong model output has a real cost.

AI/ML Consultant

SpinX AI · independent practice
  • Deliver LLM products for clients across the US, UK, Europe and Australia, owning the arc from discovery call to deployed system.
  • Translate governance requirements — AI management-system standards, national AI ethics and cybersecurity frameworks — into concrete technical controls, then ship against them.
  • Build RAG and agentic systems with evaluation harnesses and monitoring wired in from the first commit.

LangGraph · LangChain · FastAPI · EKS · Kafka · Redis · Langfuse

AI Research Engineer

Brainwave Science · contract
  • Raised real-time EEG signal detection accuracy by around 20% with a hybrid CNN/RNN architecture, benchmarked against published research and validated for reliability rather than headline accuracy.
  • Designed a modular inference pipeline with hot-swappable model components, so experiments run without disrupting production streams.
  • Turned current papers into live model iterations, then explained the results to non-technical stakeholders without the jargon.

PyTorch · TensorFlow · Keras · Python

SDE, Python & Generative AI

Asha Health · AI medical scribe
  • Built and ran a production medical scribe extracting clinical entities from physician–patient conversations, cutting manual transcription time by roughly 40%.
  • Designed LangGraph stateful agent workflows with conditional branching across documentation, coding and review stages.
  • Shipped a RAG pipeline on AWS Bedrock grounding every generated note in patient history and clinical guidelines — what made the output usable on protected health data.
  • Owned prompt governance and observability through Langfuse: versioned prompts, traced every model call, watched token spend, latency and quality to catch regressions before clinicians did.

Python · LangChain · LangGraph · AWS Bedrock/Lambda/S3/DynamoDB · Langfuse · TypeScript

SDE II, Backend & Generative AI

EMB Global · Rayo AI platform
  • Built the agentic content pipeline behind an AI SEO platform: multi-step agents for research, drafting and optimisation, with brand-voice and compliance checks enforced at generation time.
  • Cut manual intervention by roughly 60% by orchestrating content jobs, webhook triggers and third-party integrations through n8n.
  • Owned latency and scalability of the LLM serving path under high content throughput.

Python · FastAPI · LangChain · n8n · GPT-4 / Claude · Azure · PostgreSQL

Backend Developer

National Health Authority, India · C4GT open source
  • Built the subscription management microservice for the ABHA national health portal from scratch, automating consent-driven subscription flows across health services.
  • Selected for Code for GovTech, a competitive national mentorship programme for Digital Public Goods.

Spring Boot · Java · MongoDB · Docker

About

Builder, not
just a prompter

I build LLM products end to end and sit with the customer while doing it. That combination is the whole job description: most AI projects do not fail on model quality, they fail because nobody translated what the business actually needed into something a system could be held to.

My first three years were spent in places where that gap has teeth. A medical scribe running on protected health information, where an ungrounded sentence is a clinical risk. India's national health identity platform, where a consent flow has to be right for a population, not a cohort. EEG classification, where a 20% accuracy claim has to survive a benchmark against published research rather than a cherry-picked test set.

What came out of that is a fairly opinionated way of working. Ground every generation in something retrievable. Version your prompts like code. Trace every model call from day one. Put an evaluation suite between the model and your users, and let it decide what ships. Redact before you infer. None of it is exotic — it is just the difference between a demo and a system.

Today I run an independent practice, working with founders and product teams in the US, UK, Europe and Australia who need someone who can scope the problem in the room and then go and build it.

Based
Remote · India (IST)
Overlap
Full APAC & AU day · Europe afternoons · US Eastern mornings
Education
B.Tech, Electronics & Communications Engineering, IIIT Delhi (2020–2024)
Research
Undergraduate research assistant, VLSI Lab, IIIT Delhi
Certified
Deep Learning Specialization · TensorFlow Developer · Generative AI for Data Privacy & Protection
Recognition
Google Hash Code — top 900 of 2,447 teams · CodeChef 4-star · Code for GovTech
Open source
Contributor to India's national digital health infrastructure

Questions

Straight
answers

If yours is not here, ask it in the form below — I answer briefs personally.

What does an applied AI engineer actually deliver?

A working system, not a slide deck. That usually means an agent or retrieval pipeline running behind an API, wired into your data, with evaluation and tracing in place so you can see what the model did and why. I scope it with your stakeholders in plain language, build the thin slice first, instrument it, then iterate against real usage.

Do you work with regulated or sensitive data?

Yes. I spent two years on a production medical scribe operating on protected health information, and contributed to India's national digital health infrastructure through Code for GovTech.

In practice that means grounding every generated output in a retrievable source, redacting PII and PHI before it reaches a model, versioning prompts, tracing every call, and agreeing the regulatory and security constraints in writing at the start of the engagement.

What is a typical engagement and timeline?

Three shapes. A one to two week discovery sprint that ends in an architecture and a costed roadmap. A four to six week MVP build that ends in a deployed system your users can touch. Or an ongoing fractional engagement, typically two to three days a week, embedded with your team.

Which models and frameworks do you build on?

LangGraph and LangChain for orchestration, FastAPI for serving, AWS Bedrock or a direct provider API for inference, FAISS or Chroma for retrieval, and Langfuse for tracing and evaluation.

I stay provider-neutral — Anthropic, OpenAI and open-weight models all have a place, and the choice should follow the evaluation results rather than the other way round.

Can you take an existing prototype to production?

That is most of the work I get asked for. The gaps are usually the same: no evaluation harness, no tracing, prompts living in application code, retrieval that is not really grounded, and no plan for cost or latency under load. I close those, then put the serving path on infrastructure that can carry it — Kubernetes, Kafka and Redis when the load genuinely demands them.

Which time zones do you cover?

I work remotely from India on IST. That gives a full overlap with the APAC and Australian working day, a comfortable overlap with Europe, and US Eastern mornings. Clients in the US, UK, Europe and Australia have all worked in that rhythm.

How do we start?

Send a short brief through the form below, or email dsuyash57@gmail.com. If it looks like a fit, we take a 30 minute call to pressure-test the problem, and I come back with a proposed scope, a timeline and a fixed price for the first phase.

Contact

Let's build
what's next

Tell me the outcome you need and the constraint you're up against. If I'm not the right person for it, I'll say so and point you at someone who is.

This composes the message in your own email app — nothing is transmitted from this page, and no analytics or tracking runs here.

Or write to dsuyash57@gmail.com directly.