aman_hanspal
AI / ML Engineer · Mumbai, India

Aman
Hanspal

Aman Hanspal

Generative systems across image · video · audio · language

I build end-to-end AI pipelines — from research and fine-tuning to production deployment — across diffusion models, speech, retrieval, and computer vision.

aman@portfolio — ~/
➜ ~
// selected work

Experience

Aeos Labs Bengaluru
May 2024 – Jan 2026

Generative AI Researcher

  • Built generative AI pipelines across multimodal domains — image, video, speech, and automation.
  • Developed proofs-of-concept spanning RAG, speech synthesis, document intelligence, and enterprise AI workflows.
  • Translated business requirements into deployable systems with production-oriented thinking.
  • Python
  • FastAPI
  • LangChain
  • n8n
  • AWS
  • Replicate
// projects

Selected Projects

12 projects · 4 domains

Singing Voice Synthesis

featured

DiffSinger pipeline for an Ogilvy × Cadbury campaign that reached 14,000+ users at Christmas 2024. Phoneme alignment, variance modeling, and acoustic synthesis on a 3-hour custom dataset, with real-time inference on AWS Lambda.

  • DiffSinger
  • OpenUtau
  • Python
  • AWS

FaceFusion on Replicate

Deployed FaceFusion for internal and marketing use. Accepts a per-segment JSON faceswap config, applies it frame-accurately, and stitches output with ffmpeg.

  • Modal
  • Replicate
  • ffmpeg
  • Python

VisionQuery

Natural-language search over video via open-vocabulary detection. Frame sampling + a YOLO-World inference pipeline for dynamic classes with no retraining, exposed through upload/query/timestamp APIs.

  • Python
  • FastAPI
  • YOLO-World

SIGGRAPH 2025 RAG Search

Full-stack RAG over 11,000+ SIGGRAPH paper chunks. Hybrid retrieval (semantic + BM25), Cohere re-ranking, cited LLM answers, and real-time SSE streaming.

  • Next.js
  • FastAPI
  • Qdrant
  • OpenRouter
  • Cohere

Clip-It

Local-first multimodal video search for editors. Drop in footage, search in natural language, export timestamped clips — fully on-device. Dual retrieval (BGE transcript + OpenCLIP keyframe vectors) fused per video, evidence-gated chat that rejects uncited answers, and a 30-query eval harness that locks clip suggestions until top-3 accuracy ≥ 0.70.

  • Electron
  • React
  • FastAPI
  • OpenCLIP
  • Whisper
  • LanceDB
  • Ollama

RAG ChatBot · Tata Motors

Enterprise RAG for querying internal knowledge bases in natural language. Hybrid retrieval (embeddings + BM25), re-ranking, and RAGAS-based quality evaluation.

  • LangChain
  • LlamaIndex
  • RAGAS

Speech-to-Speech · JioStar

Led R&D on an English-to-Indian-regional-language speech POC. Benchmarked pipelines on latency, speed, and accuracy in a technical report; shipped on OpenAI's Realtime API.

  • OpenAI Realtime
  • Python

Document Analysis

End-to-end PDF parser for scanned and digital documents, handling mixed content — text, tables, and images — with PyMuPDF. OCR and layout extraction via Reducto and Qwen3-VL (through OpenRouter), normalized into a schema and saved as structured JSON for downstream LLM analysis. Async Celery + Redis pipeline for parallel processing with retry logic and API rate limiting.

  • Python
  • PyMuPDF
  • Reducto
  • Qwen3-VL
  • Celery
  • Redis
  • Supabase

Pinterest Scraper

Open-source Python package for scraping Pinterest images — search, download, HTML gallery generation, rate-limit handling, and a CLI. Published on PyPI.

  • Python
  • Playwright
  • PyPI

AI Films & Video

featured

Brave New Art (JioHotstar × Lenovo) — India's first fully AI-generated short film, produced end-to-end. Plus a Mother's Day film for SMFG India Credit.

  • FLUX
  • Midjourney
  • Runway
  • Kling AI
  • Hunyuan
  • ComfyUI

Image Generation

Generated images in varied artistic styles. Fine-tuned FLUX, SDXL, and SD 3.5 LoRAs on Kohya_SS and published models to CivitAI.

  • Stable Diffusion
  • FLUX
  • LoRA
  • Dreambooth
  • ComfyUI
  • A1111
// hackathons

Hackathons

2 builds

ReVision AI

Jul 2026

Agentic study companion for VideoDB's "Unlock the Footage" Global Media Intelligence Hackathon. Turns any lecture video into a full study pack — grounded clips, an AI-generated explainer video (FLUX visuals + TTS narration + auto-captions), flashcards, and summaries — from one natural-language prompt. A GPT-4o tool-loop agent picks the tools and streams live progress; optional Telegram delivery.

  • VideoDB
  • FastAPI
  • React
  • TypeScript
  • GPT-4o
  • FLUX

Note·It

May 2026

Desktop learning agent for educational videos (VideoDB Global Online Hackathon). A floating Electron overlay captures screen + audio via VideoDB, indexes streams with multimodal models, and writes structured notes into Notion.

  • Electron
  • React
  • FastAPI
  • VideoDB
  • Notion
// stack

Skills

Languages
Python · JavaScript · TypeScript · SQL · HTML · CSS
Backend & APIs
FastAPI · REST APIs · Celery · Redis
Generative AI / ML / CV
LangChain · LangGraph · LlamaIndex · FLUX · SDXL · SD 3.5 · DiffSinger · ComfyUI · LoRA fine-tuning · YOLO-World (open-vocab detection) · speaker diarization · document OCR / multimodal parsing
Frontend
React · Next.js · Vite · Electron
Cloud & DevOps
AWS · Replicate · Modal · Docker · Git · Vercel
// writing

Writing

speech ai

A Guide to Speaker Diarization

A practical breakdown of identifying who spoke when in an audio stream.

Read ↗

// reading

Reading List

Papers I read, break down, and implement.

Attention Is All You Need — 31 May '26

// education

Education

Mumbai · Class of 2024

DY Patil University, RAIT

B.Tech in Computer Engineering · CGPA 9.22

// extras

Recognition

AI Masterclass — Cohort 1

8-week intensive by Dr. Arjun Jain (Jan–Mar 2026). Certificate ↗

Patent (filed)

Enhancing organisational security using augmented-reality smart glasses · App. No. 202321037263

First Place — Anveshan 2023

Student Research Convention (West Zone) for ThirdEye — AR smart glasses for security.

// contact

Get in touch

➜ ~ echo $CONTACT Let's build something.
emailamanhanspal05@yahoo.com
phone+91 93267 66074
socialgithub · linkedin · medium