Machine learning | Computer vision | Research

Gourab Roy

Machine Learning Engineer
and Computer Vision Researcher

I am a machine learning engineer focused on visual intelligence and practical AI systems. I move between model development, applied retrieval, and the product work needed to make those systems useful.

My current interests include multimodal learning, medical image analysis, retrieval quality, and reliable LLM workflows.

Selected papers

Peer-reviewed work, accepted research, and independent study.

03

Preprint | Independent research

AstraQ-VL

An astronomy vision-language model that aligns frozen visual features with a language model, then adds instruction tuning with LoRA.

  • Vision-language models
  • LoRA
  • Astronomy
Read preprint ->

Systems I have built

GitHub profile ->

Compact CLAD

An action-focused version of CLAD that keeps visual observations and endpoint latents as context. Trained and evaluated on NIV and CrossTask, with checkpoints, logs, and five-seed results published.

  • Procedure planning
  • Diffusion models
  • PyTorch
View project ->

AstraQ-VL

A two-stage astronomy vision-language model. It aligns frozen CLIP ViT-L/14 features with Qwen2.5-1.5B-Instruct, then adds LoRA instruction tuning for visual question answering on astronomical imagery.

  • Multimodal learning
  • CLIP
  • Qwen
  • LoRA
View project ->

TerraQ-VL

A vision-language model for aerial and satellite imagery. It adapts the AstraQ-VL training recipe to VRSBench with a frozen CLIP encoder, Qwen2.5-3B, connector alignment, and LoRA instruction tuning.

  • Remote sensing
  • CLIP
  • Qwen
  • LoRA
View project ->

DocuMindGPT

A document-grounded question-answering CLI with retrieval-augmented generation, support for PDFs and text files, and an included evaluation step for relevance and hallucinations.

  • Language systems
  • RAG
  • Supabase
  • Gemini
View project ->

CT Image Reconstruction

An end-to-end deep learning pipeline that reconstructs CT images from sinograms, developed for sparse-view and low-dose acquisition scenarios.

  • Medical imaging
  • PyTorch
  • Swin Transformer
View project ->

LLM Chat Agent

A conversational AI application with query routing, sub-query decomposition, NDJSON streaming, resumable sessions, and optional web grounding.

  • LangGraph
  • FastAPI
  • React
View project ->

Work and education

  1. 2025 - Present

    Data Scientist

    Axtria, Bengaluru

    Working on multi-agent LLM workflows, retrieval-augmented analytics, and real-time conversational systems for enterprise use.

  2. May 2025 - Sept 2025

    Research Intern, Deep Learning for Medical Imaging

    Indian Institute of Technology, Kharagpur

    Built a modified U-Net for CT image reconstruction on 20,000+ sinogram-image pairs, reaching PSNR above 35 dB and SSIM above 0.9 for sparse-view and low-dose scenarios.

  3. May 2024 - July 2024

    Data Science Intern

    Axtria, Bengaluru

    Built and validated time-series forecasting models and automated recurring analytics for planning work.

  4. 2023 - 2025

    M.Tech, Computer Science and Engineering

    IIT (ISM) Dhanbad

    Thesis on a multi-stage deep learning framework for knee osteoarthritis and osteoporosis detection, producing two peer-reviewed papers including a Best Paper Award at ISAI 2025.

  5. 2019 - 2023

    B.Tech, Electronics and Telecommunications Engineering

    IIEST Shibpur

    Coursework included signal processing, digital systems, programming, and linear algebra.

ML blog

Thoughts and writing on machine learning.

The archive covers topics such as meta-stacking, adversarial validation, model evaluation, deep learning, and optimization.

Browse articles ->

04 | Contact

Open to research collaborations and ML opportunities.