I am a Senior Quant Research Engineer bridging the gap between High-Performance Computing and Alpha Generation.
- π Currently building: Autonomous DeFAI trading agents (LLM + RL) on distributed infrastructure.
- π Education: MS CS (Machine Learning) @ Georgia Tech | CQF (Quantitative Finance).
- β‘ Core Stack: Python, C++, PyTorch, CUDA, Slurm, AWS ParallelCluster, Kubernetes, Docker.
- π Focus: HFT Infrastructure, Market Microstructure, and AI-driven Trading Strategies.
Website | LinkedIn | X / Twitter
slurm-scheduler-lab β test Slurm priority and backfill policy against a job trace before it reaches a live controller.
Implements priority/multifactor and EASY backfill, reads PriorityWeight* straight from a slurm.conf, and replays real sacct traces. I built it to answer a question I couldn't safely test in production β and the answer turned out not to be the weights:
backfill OFF backfill ON
cpu utilization 72.2 % cpu utilization 83.6 %
mean wait 1913.0 min mean wait 373.7 min
More in progress β agentic inference infrastructure, 128-GPU distributed training, and CUDA volatility surface calibration. They go public as they get good enough to defend.
I publish at zhanyl-tech.github.io β deep dives on inference optimization and HPC, plus shorter lab notes on whatever I'm currently measuring.
- Your Slurm priority weights matter less than your users' time limits
- LServe and SampleAttention: what sparse attention actually changes
- How KV-cache paging works in vLLM
βοΈ Chess and poker outside of work β both cheaper places to practise reasoning under uncertainty than production is.