MLOps · DevOps · Cloud · AI/ML · Data Engineer
Bridging the gap between ML research and robust production systems
I'm a passionate MLOps, DevOps, Cloud & AI/ML Engineer with a deep focus on building scalable machine learning systems that actually ship to production. I specialize in designing end-to-end ML pipelines, automating infrastructure, and integrating large language models into real-world applications.
From containerizing models with Docker and orchestrating them on Kubernetes, to fine-tuning LLMs and building RAG pipelines — I work across the full stack of modern AI infrastructure. My cloud experience spans AWS, GCP, and Azure, and I'm deeply familiar with tools like MLflow, Kubeflow, Airflow, Terraform, and the broader LangChain ecosystem.
I believe in open-source collaboration, continuous learning, and building systems that are reproducible, monitored, and maintainable. Currently contributing to trending OSS projects and always open to exciting challenges at the intersection of AI and infrastructure.
A diverse skill set spanning the full AI/ML lifecycle and modern cloud infrastructure
My own projects first — each one tested, CI-green and runnable in one command — followed by the open-source ecosystem I build on and contribute to (stars shown belong to the upstream projects)
One API in front of many LLM providers: cheapest-first routing, response caching, automatic failover with circuit breaking, and a cost ledger that separates spend from spend avoided. Runs offline — clone and run in one command.
Golden-set evaluation for RAG systems as a deterministic CI gate — a change that degrades answer quality fails the build and names the case that broke. Catches the fluent answer with the load-bearing fact missing.
Reads NVIDIA DCGM metrics and re-exports what utilisation graphs don't show: burn, waste (the idle share, in money) and cost per 1k inferences, as Prometheus metrics. Stdlib only, ships a captured DCGM sample.
Get up and running with large language models locally. Simplifies LLM deployment on personal hardware with a clean REST API and model management.
High-throughput and memory-efficient inference engine for LLMs. Uses PagedAttention for 24x faster serving vs HuggingFace. PyTorch ecosystem member.
Production-ready LLM app development platform with RAG pipelines, agentic workflows, model management, and observability in one intuitive interface.
End-to-end ML platform with experiment tracking, model registry, LLM observability, and AI agent tracing. 30M+ monthly downloads. The industry standard.
ML Toolkit for Kubernetes enabling end-to-end ML workflows. Covers pipelines, Katib hyperparameter tuning, and model training at scale. CNCF member.
The de facto standard for building LLM-powered applications. Modular abstractions for chains, agents, RAG, memory, and tool use. Massive plugin ecosystem.
AI compute engine providing a distributed runtime and ML libraries for accelerating ML workloads. Used by Uber, Spotify, OpenAI. 27M monthly downloads.
Platform to programmatically author, schedule, and monitor workflows. The backbone of data engineering pipelines at thousands of companies worldwide.
Workflow orchestration for building resilient, dynamic data pipelines in Python. Reacts to real-world events, Pythonic API, Kubernetes-native.
State-of-the-art ML models for text, vision, audio, and multimodal tasks. 1.2B+ installs, 750K+ models, Transformers v5 with modular architecture.
IBM's open-source SDK for quantum computing. Compose and run quantum circuits on real IBM Quantum hardware or local simulators. Actively exploring hybrid classical-quantum ML algorithms for optimization and sampling problems.
Low-code workflow automation platform wired to Claude API for end-to-end AI agent pipelines. Build multi-step agentic workflows — web research, RAG retrieval, email drafting, Slack alerts — all orchestrated via n8n's visual node graph.
Staying current with the latest breakthroughs, model releases & community updates
Meta’s Muse Spark achieves Llama 4 performance at an order of magnitude less compute, built by Meta Superintelligence Labs with $115-135B capex backing.
Read more →The first widely recognized 10-trillion-parameter model will not receive public release due to cybersecurity risks, rolling out to a handpicked consortium only.
Read more →OpenAI’s GPT-5.4 Thinking variant scores 75% on OSWorld-Verified, officially surpassing human-level performance on desktop task benchmarks.
Read more →Grok 4.20 features a 4-agent system (coordinator, research, logic, contrarian) working in parallel and cross-verifying outputs for unprecedented reliability.
Read more →Record-breaking VC quarter dominated by OpenAI, Anthropic, and major acquisitions. Over 30 new models launched in March alone.
Read more →New compression algorithm maintains frontier performance while cutting memory requirements by a factor of six — a game-changer for edge AI deployment.
Read more →Curated AI/ML tutorials updated weekly by AI agent
TechWorld with Nana
TechWorld with Nana
Vishakha Sadhwani
CNCF / Lin Sun
IBM Technology
Suresh Raju
Active in the MLOps and AI engineering community — connecting, contributing, and sharing knowledge
Exploring and sharing models, datasets, and spaces. Following cutting-edge model releases and community discussions.
Visit ProfileActive open-source contributor. Forking, PR-ing, and starring repos across the MLOps, LLM, and cloud-native ecosystems.
GitHub ProfileSharing insights on MLOps, AI engineering, and cloud infrastructure. Open to professional connections and collaborations.
ConnectFollowing AI researchers, MLOps practitioners, and sharing quick insights on tools, papers, and engineering patterns.
FollowWriting about MLOps, DevOps, and AI engineering. Sharing practical insights on Kubernetes, LLMs, CI/CD, and cloud-native infrastructure.
Read ArticlesMember of the global MLOps community — attending virtual meetups, reading the blog, and participating in discussions.
Join CommunityWriting tutorials and deep-dives on MLOps patterns, LLM deployment strategies, and cloud infrastructure best practices.
Read ArticlesConsistent open-source activity across AI, MLOps & DevOps repositories
Open to full-time roles, freelance projects, and open-source collaborations
Whether you're looking for an MLOps engineer to scale your ML infrastructure, a DevOps expert to streamline your CI/CD, or just want to chat about the latest in AI — my inbox is open!
I typically respond within 24 hours.