PhD candidate · Pattern Recognition Lab, FAU

Hi, I'm Md Hasan

Teaching machines to read images — and write about them.

I research AI for medical imaging at the Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg — where computer vision meets large language models. Before the PhD I spent five years building production ML: face recognition and anti-spoofing systems at Gaze and CPSD Technologies in Dhaka, then data and LLM work at the BMW Group and Innomotics (Siemens) in Bavaria.

Portrait of Md Hasan Multimodal AI Nürnberg, DE
5
Publications
& preprints
5
Years of industry
ML & data work
2
Degrees in CS
& Artificial Intelligence
PhD
In progress at the
Pattern Recognition Lab
What I work on

Research & focus areas

My thread through all of it: getting models to reason across more than one modality at a time, and making them dependable enough to leave the notebook.

Multimodal AI & medical imaging

Vision–language models that read a scan and produce clinically faithful text. My MSc thesis generated region-guided clinical reports from chest X-rays using LLMs — the PhD at the Pattern Recognition Lab carries that line of work forward.

Vision–Language Chest X-ray Report Generation

LLMs, RAG & generative AI

Retrieval-augmented systems that stay grounded in their sources. At Innomotics I built RAG applications over dense industrial documentation — evaluation, chunking strategy and hallucination control included.

RAG Fine-tuning Evaluation

Computer vision at the edge

Face anti-spoofing, detection and landmark models small enough to run on a phone. Three of my papers come from this area, and I shipped it for real — a facial attendance system in production, 3× faster on edge devices.

Anti-spoofing ONNX / ncnn Detection
Toolbox

Things I build with

Languages

PythonC++SQLBash

Deep learning

PyTorchTensorFlowKerasHugging FaceOpenCVscikit-learn

Generative AI

LLMsRAGPrompt EngineeringFine-tuningVector search

Edge & optimisation

ONNXncnnKnowledge DistillationPruningQuantisation

Data & MLOps

DockerMLOpsETL / ELTPandasNumPySQLAlchemyPlotly

Languages spoken

Bangla — nativeEnglish — professionalGerman — elementary
The path so far

Experience & education

Jul 2025 — Present

PhD Candidate, Pattern Recognition Lab

FAU Erlangen-Nürnberg · Erlangen, Germany
Researching multimodal deep learning for medical image understanding — vision–language models that turn imaging studies into accurate, clinically useful text.
Apr 2022 — Jun 2025

MSc, Artificial Intelligence

FAU Erlangen-Nürnberg · Erlangen, Germany
Thesis: Generation of region-guided clinical text reports from chest X-ray images using LLMs, supervised by Tomas Arias-Vergara, Paula Andrea Pérez-Toro and Andreas Maier.
Oct 2023 — Mar 2025

Data Scientist (Working Student)

Innomotics GmbH — A Siemens Business · Nürnberg, Germany
  • Built a configurable, reusable internal RAG framework powered by LLMs.
  • Designed chunking, embedding and evaluation pipelines to keep answers grounded in source documents.
  • Packaged prototypes into deployable services for non-technical business users.
Sep 2022 — Aug 2023

Data Analytics (Working Student)

BMW Group · Munich, Germany
  • Developed ETL workflows feeding analytics used across internal business processes.
  • Built predictive models and dashboards that replaced manual reporting.
Mar 2021 — Mar 2022

AI Engineer

CPSD Technologies Ltd. · Dhaka, Bangladesh
  • Engineered the face recognition framework behind HAJIRA, an automated facial attendance system used by multiple companies.
  • Made face detection and recognition 3× faster on edge devices, from video and live streams.
  • Cut inference time by 50% through optimisation and an efficient ncnn implementation.
  • Built an optimisation framework — simplification and pruning — to deploy pretrained models on low-powered devices.
Feb 2020 — Jan 2021

Deep Learning Engineer

Gaze · Dhaka, Bangladesh
  • Pushed face liveliness detection to 99% accuracy on an in-house dataset, shipped inside GazePass, a passwordless authentication system.
  • Built a synthetic NID image dataset with conditional GANs for rigorous evaluation on regional data.
  • Shrank the face recognition pipeline using knowledge distillation, with no loss of validity.
Jan 2016 — Apr 2020

BSc, Computer Science & Engineering

North South University · Dhaka, Bangladesh
Thesis: BanglaKotha — Bangla automatic speech recognition leveraging RNN-T. Champion of Bangladesh's first university-level Datathon (NSU / ACM, 2019), and first author on work in face anti-spoofing and Bangla NLP.
Selected work

Projects

A few things I've built end to end, from research code to deployable artefacts.

Face detection with landmark keypoints
2022 · Computer Vision

Face Detector with Landmarks

A sub-1 MB face detector with 5-point landmarks, built for mobile and edge deployment across ncnn, ONNX and PyTorch.

PyTorchONNXncnn
Read more
Electric vehicle charging analysis
2023 · Data Engineering

EV Charging Patterns in Germany

An end-to-end ELT pipeline and analysis of Germany's EV charging infrastructure — regional gaps, capacity and where to build next.

PythonSQLPlotly
Read more
Bangla speech recognition waveform
2020 · Speech

Bangla Speech Recognition

A sequence-to-sequence ASR system for Bangla in PyTorch, with the full data, training and validation toolchain around it.

PyTorchSeq2SeqLibrosa
Read more
Peer reviewed

Recent publications

2026

A Deep Risk Estimator for Known Operator Learning

Andreas Maier, Md Hasan, Paulina Conrad, Paula Andrea Perez-Toro

arXiv preprint · under review

2022

MHASAN: Multi-Head Angular Self Attention Network for Spoof Detection

Md Hasan, Koushik Roy, Labiba Rupty, Md. Sourave Hossain, Shirshajit Sengupta, Shehzad Noor Taus, Nabeel Mohammed

ICPR 2022 — 26th International Conference on Pattern Recognition

2021

Bi-FPNFAS: Bi-Directional Feature Pyramid Network for Pixel-Wise Face Anti-Spoofing by Leveraging Fourier Spectra

Koushik Roy, Md Hasan, Labiba Rupty, Md. Sourave Hossain, Shirshajit Sengupta, Shehzad Noor Taus, Nabeel Mohammed

Sensors (MDPI) · 21(8), 2799

All publications with abstracts

Let's talk research — or anything multimodal

I'm always up for conversations about vision–language models, medical AI, collaborations, or a good supervision question. My inbox is open.