YuZhou Wu
Hi! I am YuZhou Wu, a Master’s student in Robotics at the University of Sheffield (expected Sept 2025).
My research interests lie at the intersection of Robotics, Vision-Language-Action (VLA) models, and efficient spatiotemporal representation learning, with a particular focus on 4D perception, multi-agent reasoning, and Transformer acceleration for real-time embodied systems.
I am actively preparing for PhD applications and am interested in research-oriented positions related to robot learning, embodied AI, and large-scale multimodal models.
🔍 Research Interests
My current research focuses on:
- Vision-Language-Action (VLA) Models
- Policy learning for embodied agents (e.g., LIBERO, DROID-style datasets)
- Action tokenization, latent action modeling, and policy generalization
- Streaming inference and real-time deployment constraints
- 4D Point Cloud Perception & Transformers
- PointTransformer / SpUNet architectures for dynamic scenes
- Cross-frame attention and KV-Cache enhanced Transformers
- World-coordinate alignment and temporal token reuse
- Efficient & Interpretable Deep Models
- Token pruning, attention sparsification, and quantization-aware training
- KV cache reuse for streaming perception
- Model compression for deployment on resource-constrained systems
🧪 Selected Research Projects
KV-PT: Key-Value Cache Enhanced Point Transformer for 4D Space Analysis
- Proposed a temporal KV-cache mechanism for PointTransformer-style models
- Enables cross-frame attention with nearest-neighbor and max-attention token reuse
- Applied to SemanticKITTI sequential LiDAR data
IPCV: Information-Preserving Compression for Vision Encoders
- Proposed an information-preserving compression framework for vision encoders, enabling token reduction without sacrificing semantic fidelity
- Designed structure-aware token pruning and feature reweighting mechanisms, guided by representation consistency constraints
- Achieved strong performance–efficiency trade-offs across multiple vision backbones, improving inference speed and memory usage
Multi-Agent Motion Prediction & Decision-Making Survey
- Systematic taxonomy of multi-agent prediction and planning methods
- Comparative analysis on Waymo Open Motion & Argoverse2
- Covers game-theoretic, MARL, intent-sharing, and V2X-based approaches
🛠 Technical Skills
- Programming: Python, C++, MATLAB
- Deep Learning: PyTorch, TorchScript, ONNX
- Robotics: ROS2, Nav2, SLAM, TurtleBot3
- 3D / 4D Vision: PointTransformer, SpConv, SemanticKITTI
- Tools: Git, Docker, Conda, VSCode Remote, tmux
📌 Current Goals
- Submitting papers to CVPR / ECCV / NeurIPS / ACL
- Applying for PhD programs in Robotics / Embodied AI
- Exploring efficient VLA models for real-time robotic control
Feel free to explore my publications, projects, and code repositories, or reach out if you are interested in collaboration.
