DDOOVVIISSBBLLOOGG

To be the most intuitive, most powerful, most unified engineer.
TechnologyStackandFeaturedProjects
Explore some of my latest work and creative endeavors
  • React
  • Next.js
  • TypeScript
  • Tailwind CSS
  • Python
  • TensorFlow
  • PyTorch
  • Java
  • Spring
  • Apache Kafka
  • Node.js
  • GraphQL
  • Docker
  • Kubernetes
  • MongoDB
  • PostgreSQL
  • Redis

Early Exit LLM Research

Advanced KV cache and Page-attention learning for memory optimization.

GPU pipelinePytorchAdaptively RL

GNNs Hypergraph Research

Exploring graph neural networks for hypergraph data representation.

GNNsGraph TheoryGANs

MiMo-Audio on vLLM Omni

Async audio-token streaming and CUDA Graph capture enable low-latency multimodal inference.

vLLM OmniCUDA GraphsAudio Tokens

Quantum Computing Simulator

Scalable data processing and analytics platform for enterprise stocks solutions.

RL StrategyKubernetesApache Spark

LLM Quantification Training

PTQ and RTQ quantization aware training for large language models.

PTQDeepSpeedRTQ

CUDA Programming

High-performance parallel computing with CUDA for deep learning applications.

CUDAGPU OperatorParallel Computing

Qwen3-VL EAGLE-3

Draft-model training accelerates multimodal inference without sacrificing quality at production scale.

EAGLE-3Qwen3-VLDeepSpeed ZeRO-2

Distributed Draft Training

DeepSpeed ZeRO-2 and Accelerate train production-aligned draft models across 64 H20 GPUs.

DeepSpeed ZeRO-2Accelerate64× H20

Early Exit LLM Research

Advanced KV cache and Page-attention learning for memory optimization.

GPU pipelinePytorchAdaptively RL

GNNs Hypergraph Research

Exploring graph neural networks for hypergraph data representation.

GNNsGraph TheoryGANs

MiMo-Audio on vLLM Omni

Async audio-token streaming and CUDA Graph capture enable low-latency multimodal inference.

vLLM OmniCUDA GraphsAudio Tokens

Quantum Computing Simulator

Scalable data processing and analytics platform for enterprise stocks solutions.

RL StrategyKubernetesApache Spark

LLM Quantification Training

PTQ and RTQ quantization aware training for large language models.

PTQDeepSpeedRTQ

CUDA Programming

High-performance parallel computing with CUDA for deep learning applications.

CUDAGPU OperatorParallel Computing

Qwen3-VL EAGLE-3

Draft-model training accelerates multimodal inference without sacrificing quality at production scale.

EAGLE-3Qwen3-VLDeepSpeed ZeRO-2

Distributed Draft Training

DeepSpeed ZeRO-2 and Accelerate train production-aligned draft models across 64 H20 GPUs.

DeepSpeed ZeRO-2Accelerate64× H20

Ready to explore?

Explore the systems behind the benchmarks—from multimodal audio serving to speculative decoding and adaptive edge inference.

Spacer for better visual balance