Mohammad Wasil Saleem
I'm a Machine Learning Engineer
Professional Summary
Machine Learning Research Engineer with 4 years of experience specializing in 3D Computer Vision, Neural Rendering, NLP, and Generative AI. My work involves dedicated research in Deep Learning, GenAI, Multimodal AI, and Reinforcement Learning, coupled with architecting production-scale MLOps workflows. My portfolio further explores the development of Games and 3D modeling.
Work Experience
Machine Learning Research Engineer - 3D Computer Vision & Neural Rendering
- Neural Rendering: Trained 3D Gaussian Splatting (3DGS) and NeRF representations to perform photorealistic novel view synthesis from arbitrary camera positions, bypassing physical data collection bottlenecks and saving €10k+ per capture event.
- To eliminate under-observed blind spots and volume artifacts, I rendered novel camera trajectories from base incomplete Gaussians, used generated video captions and camera conditioning to generate enhanced views via a conditioned video diffusion model, and distilled the inpainted frames back into a second Gaussian optimization stage.
- Trained depth estimation architectures, introducing consistency-aware loss functions, alongside train and test-time augmentations to stress-test robustness and improve depth accuracy.
- Orchestrated Structure-from-Motion (SfM) and point cloud registration workflows, utilized COLMAP for precise camera pose estimation and Open3D for depth fusion.
- Converted 3D point clouds into high-fidelity meshes using Marching Cubes, Poisson Surface Reconstruction, and Nvblox, exporting metric-scale assets into Isaac Sim for simulation.
- Verified point cloud alignment and 3D reconstruction fidelity using CloudCompare, Blender, COLMAP, and Isaac Sim.
- Engineered robotic simulation environments in Isaac Sim by applying spatial transformations, dynamic rigid-body physics, and animated human agent trajectories, generating occupancy maps and automated multi-sensor (RGB-D) data collection pipelines for model training.
- Extracted 3D human Gaussians from images using NVIDIA Asset Harvester for simulation.
- Tech Stack: Python, PyTorch, Git, Bitbucket, Linux, Isaac Sim, Open3D, COLMAP, CloudCompare, Blender, SuperSplats, Nerfstudio
Machine Learning Software Engineer
- Siemens AI for Engineering : Project focused on the development and integration of AI solutions to automate domain-specific engineering workflows and significantly boost efficiency.
- Artificial intelligence: transforming mobility for everyone: Company AI page
- Designed and implemented Quality Agent, a company-wide tool for requirements analysis.
- Developed a scalable multi-AI agent framework for automated test case generation for train braking systems.
- Deployed across 7 distinct use cases, analyzing around 20 requirements per day.
- Developed a self-service portal in JavaScript (GUI) and SQLite cache to improve adoption.
- Integrated seamlessly with requirement management systems by developing secure APIs.
- Engineered an automated Table Extraction pipeline (OCR) utilizing cloud-based AI solutions.
- Designed to efficiently process highly unstructured, scanned technical documents.
- Tech Stack: Python (LangGraph, Pydantic), JavaScript, SQL, Docker, Linux, AWS (IAM, S3, Bedrock, Textract), Azure (OpenAI, Document Intelligence), Git, CI/CD, Poetry, mypy, pre-commit.
Machine Learning Research Engineer - Computer Vision (Autonomous Trains)
- Siemens Driverless Train : Project focused on Assistance and driverless train operations to maximize system capacity, improve safety, and ensure sustainability.
- Assisted and driverless train operation: Company page
- Developed and fine-tuned an end-to-end pedestrian detection model for autonomous trains using panoptic segmentation, aligning custom dataset with COCO dataset format.
- Designed and implemented an Active Learning framework (uncertainty and diversity-based query strategies), boosting model accuracy by 89.94% across five cycles using just 7% of queried data, supporting ADAS safety.
- Self-supervised Learning: Improved clustering of unlabeled image data by fine-tuning embeddings of ResNet-152 and VGG16 pre-trained with ImageNet dataset using contrastive learning and visualizing the results using t-SNE.
- Built a semi-automated Data annotation pipeline using CVAT and models such as YOLO, SAM, MaskDINO, Mask2Former, and PanopticFCN, reducing the annotation time by 76%.
- Leveraged Multimodal foundation models (BLIP and BLIP-2) for zero-shot transfer and semantic image retrieval to streamline the data curation process across large visual datasets.
- Tech Stack: Python, PyTorch, Detectron2, CVAT, Git, Linux (WSL2).
Data Analyst - Research Assistant
- ATB Animal Welfare : Conducted statistical analysis of air exchange rates in naturally ventilated barns using methods such as ANOVA and model selection (AIC/BIC).
- Individualized Livestock Production: Company page
- Enhanced regression model performance by increasing R² from 0.65 to 0.85 through targeted feature transformation and optimization.
- Developed scripts for task automation with Python and R; documented results in RMarkdown.
- Tech Stack: R, Python, Excel, RMarkdown; Statistical and Probabilistic Modeling.
Unity3D Developer Intern
- Designed and developed the core mechanics and user interface for the Android 3D game.
- Implemented dynamic runtime 3D mesh generation for tunnels, significantly boosting in-game performance.
- Tech Stack: Unity3D, C#, JavaScript, Blender.
Education
Master of Science: Data Science
- Grade: 1.8 GPA
- Master's Thesis Topic: "Cost-Efficient and Model-Guided Multi-Query Strategy for Pedestrian Segmentation for ADAS in Railways"
- Master's Thesis Grade: 1.1 GPA
- In collaboration with Siemens Mobility
- Thesis Link: Research Gate Link for Master Thesis
- Proposal: Proposed an Active Learning-based Panoptic Segmentation method for railway safety, achieving an 89.94% AP improvement using only 7% of data points, while reducing labeling time by 76.19% through a deep learning–assisted human-in-the-loop annotation process.
Bachelor of Technology: Computer Science and Engineering
- Grade: 1.6 GPA
- Bachelor's Final Project Topic: "Visual Question Answering"
- Bachelor's Final Project Score: 463/500 (92.6%)
Skills
Languages
Programming Languages
Machine / Deep Learning Expertise / Computer Vision / LLMs
3D Reconstruction & Neural Rendering
Python Libraries & Frameworks
Engines, Simulators & Asset Tools
MLOps / DevOps
Cloud
Blog Posts
Grad-CAM: What Does a Self-Driving CNN See Under Domain Shift? Blog
Portfolio
A comprehensive portfolio featuring Machine Learning and Deep Learning research, Reinforcement Learning, Neural Rendering and GenAI, supported by Docker-based engineering and Game Dev projects.
- All
- Machine/Deep Learning
- Computer Vision / Neural Rendering
- Natural Language Processing
- GenAI
- Reinforcement Learning
- Game Development
Pedetrians Detection for Trains & Trams
Self-Driving Car (with Grad-CAM Analysis under Domain Shift)
Trained a Deep Learning-based autonomous driving model using CNNs for steering prediction and road line detection on Synthetic Data generated via Unity3D simulation, with real-time image streaming over TCP/IP. Used Grad-CAM to analyze how the model’s attention changes under different weather conditions, highlighting the effects of domain shift.
Neural Rendering for Autonomous Driving (WiP)
Neural Machine Translation with Attention
Blender Text-to-Node ChatBot (WiP)
Conditional GANs - Generate New Faces
Predicting User Affiliation of YouTube Commenters with Hierarchical Attention
VQA: Visual Question Answering (Vision + Language)
Ticker Checker
Q-learning vs SARSA applied to Smart Cab
Contact
Reach out to me!
wasilmohd1@gmail.com
Location
Potsdam, Germany