01 About

About

Final-semester Master of Data Science student at Macquarie University, Sydney, working in machine learning, deep learning, graph neural networks and statistical learning.

Published research applies convolutional regression and segmentation architectures to fetal biometry from ultrasound, alongside a second paper on metaheuristic optimisation for constrained problems. Current work extends this into graph neural networks for fraud detection and time series forecasting, alongside weakly-supervised and attention-based learning, unstructured text and image mining, and predictive modelling on real, messy, imbalanced data. Rigour matters as much as results: honest evaluation metrics, leakage-safe design, and reproducible pipelines.

Research interests: Graph Machine Learning · Deep Learning · Fraud & Anomaly Detection · Time-Series Forecasting · Weakly-Supervised Learning · Medical Imaging & Computer Vision · Image Processing

02 Publications

Peer-reviewed research

2
Peer-reviewed IEEE publications

Live citation metrics via Google Scholar (linked in sidebar).

03 Education

Education

Master of Data Science

Macquarie University, Sydney

February 2025 – December 2026 (expected)

Core focus
Graph neural networks & fraud detection, time series forecasting, statistical & machine learning methods, mining unstructured data
Relevant coursework
Statistical & machine learning methods, statistical inference, statistical computing, big data technologies, mining unstructured data, database systems

Bachelor of Science in Electrical and Electronic Engineering

Khulna University of Engineering and Technology (KUET), Bangladesh

2018 – 2022

04 Projects

Selected work

Current Research
Current · Major Project

Benchmarking Graph-Based Fraud Detection: A Unified Evaluation of State-of-the-Art Methods

Python · PyTorch · GNNs

Team project under Dr Venus Haghighi benchmarking graph-based fraud-detection methods under a common, reproducible evaluation framework, focused on class imbalance, heterophily and fraud camouflage rather than single-run leaderboard results. My individual contribution: the common evaluation pipeline, imbalance-aware metrics, multi-seed aggregation, statistical result summaries, and runtime/peak-memory profiling for performance-cost analysis.

Current · Coursework Project

Replicating and Generalising LTSF-Linear

Python · DLinear

Reproduces and extends DLinear/LTSF-Linear (Zeng et al., "Are Transformers Effective for Time Series Forecasting?", AAAI 2023) on public and controlled synthetic time series to test whether simple linear forecasters remain competitive with Transformer-based models under trend, seasonality, regime change, noise and level shifts.

Multi-Instance Learning for Medical Image Analysis

weak supervision · attention · survey

A literature survey tracing Multi-Instance Learning from classical instance- and bag-level methods to attention-based, correlation-aware and hierarchical deep models, using whole-slide pathology as the anchoring application, where a slide is a bag of patches and only slide-level labels exist. Extends my previous medical-imaging research into weakly-supervised learning, where only coarse slide-level supervision may be available.

Selected Applied Research

SafeSignal: Early Product-Safety Risk Detection

RoBERTa · PU learning · Sentence-BERT · FAISS

A prototype unstructured-text mining pipeline exploring how emerging product-safety hazards could be surfaced from customer reviews and complaint narratives. The design explores transformer hazard extraction, positive-unlabelled learning for scarce labels, semantic recall retrieval, temporal burst detection, heterogeneous graph aggregation, and evidence-grounded LLM summarisation, aiming for a ranked analyst alert with traceable evidence rather than an automated decision.

Mortgage Default Risk Modelling

R · tidymodels · imbalanced classification

Predicted serious Fannie Mae mortgage default using only origination-time credit, debt and equity signals, deliberately excluding post-origination variables that would leak the outcome. Compared logistic regression, LDA, regularised logistic regression, decision tree and random forest under leakage-safe, rare-event-aware evaluation, examining discrimination and interpretability trade-offs (ROC AUC, PR AUC, top-decile capture) for imbalanced financial-risk prediction.

Pedestrian Crash Severity Analysis

R · JAGS · MCMC

Modelled injury severity in South Australian pedestrian crashes (2019–2023) with a Bayesian logistic regression, reporting population-standardised probabilities and odds ratios. Included MCMC convergence diagnostics, posterior-predictive checks and prior-sensitivity analysis so the conclusions could be stress-tested rather than taken on trust.

Tools & Other Builds

FPL Prediction Site

data-driven · web

A Fantasy Premier League analytics build pulling player and fixture data from the official FPL API into transfer and captaincy recommendations, published under the PotFPL name.

↗ View repository

Human Activity Recognition from Smartphone Sensors

Python · scikit-learn · Random Forest

Built a supervised classifier for six activities of daily living on the UCI HAR benchmark (10,299 samples, 30 participants, 561 engineered accelerometer and gyroscope features). Random Forest was selected over SVM for robustness to high-dimensional noisy features, reaching 92.6% accuracy and 0.92 macro F1 under subject-independent evaluation, with per-class confusion analysis of the remaining errors.

Payroll & Revenue Forecasting

Python · scikit-learn · bootstrap

Forecasted a company's payroll and revenue trajectory, comparing ordinary least squares against Huber robust regression after identifying a large outlier month via IQR analysis. Justified MAE as the business-facing metric, quantified uncertainty with 1,000-iteration bootstrap confidence intervals, and critically assessed where long-horizon linear projections stop being credible.

Medical Insurance Cost Prediction

R · glmnet · Random Forest · 10-fold CV

Modelled private health-insurance claim expenses under heavy zero-inflation using a two-part hurdle framing, separating claim occurrence from claim severity. Tuned an elastic net against a random forest via stratified 10-fold cross-validation, with log-scale outcomes for comparable RMSE.

Also: R package development with unit testing and Quarto vignettes; a large-scale survey analysis (n = 2,000) using hypothesis testing and multiple regression; spatial query processing with R-tree indexing; and relational database design with stored procedures and transaction control.

05 Skills

Technical toolkit

ML / Deep learning
Graph neural networks PyTorch CNNs Image segmentation Transformers Weak supervision
Statistical learning
Regularised regression Random forests Cross-validation Imbalanced classification ROC / PR AUC Robust regression Bayesian inference Time series forecasting
Research evaluation
Reproducible evaluation Multi-seed experimentation Class-imbalance evaluation Calibration / uncertainty Statistical model comparison Leakage-safe validation
Languages & tools
Python scikit-learn pandas matplotlib R tidyverse tidymodels Quarto testthat SQL Flask FastAPI PostgreSQL LLM APIs Git
06 Experience

Professional background

AI Operations Manager Dec 2023 – Jan 2025
Grabsoft Digital · Dhaka

Led B2B client engagements integrating AI capabilities into existing software products, managing solution design and implementation from initial scoping through to delivery.

Additional training: Fundamentals of Visualization with Tableau, University of California, Davis (Jul 2025).

Last updated: October 2026