Machine learning engineer with 4 years taking computer-vision models from research to production on GPUs. Optimises inference for latency and cost without losing accuracy, and builds the data loops that keep models improving after launch. Published at a CVPR workshop.
- Deployed a defect-detection model to 40 factory lines, catching 93% of defects versus 71% by manual inspection.
- Cut inference latency from 48 ms to 9 ms with TensorRT quantisation, keeping accuracy within 0.3%.
- Built a labelling and retraining loop that improved recall 6 points per quarter.
- Reduced GPU serving costs 45% by batching requests on Triton and right-sizing instances.
- Wrote the team's model-release checklist, cutting failed deployments from 5 to 0 per quarter.
- Partnered with factory engineers to set alert thresholds that cut false alarms by half.
- Profiled training jobs and removed data-loading bottlenecks, raising GPU utilisation from 55% to 90%.
- Co-authored a CVPR workshop paper on low-light object detection.
- Trained segmentation models on 8 GPUs with mixed precision, halving training time.
- Published an open-source dataset loader now used by 3 other research labs.
- Taught weekly recitations for a 120-student computer vision course.
- Built a synthetic-data pipeline that tripled training examples for rare defect types.
- Presented the lab's work at 2 industry sponsor reviews.
- Built image-processing services in C++ handling 5M photos a day.
- Accelerated 4 image filters 12× by rewriting them in CUDA for GPU servers.
- Improved face-detection accuracy 8% by retraining on a cleaned dataset.
- Deployed a photo-quality classifier that filtered 20% of low-quality uploads.
- Wrote internal documentation for the GPU build pipeline used by 15 engineers.
Mandarin (Native) · English (Fluent)