Research interests center on LLM post-training and sequential decision-making, with current work on WebAgent systems and card-game AI. Core methods include supervised fine-tuning, reinforcement learning, self-play, and Monte Carlo tree search in imperfect-information games; earlier work focused on diffusion-based generative modeling and distributed training.
| LLM Post-Training |
|
| Agents |
|
| Game AI |
|
| Generative Modeling |
|
| Research Code |
|
| Publication | Venue | Research topic |
|---|---|---|
| SMRL · Paper | IEEE Transactions on Neural Networks and Learning Systems, 2026 | Spatial meta-learning for unseen geographic entities |
| STMetaT · Paper | Knowledge-Based Systems, 2025 | Spatio-temporal meta-learning for trajectory representation |
| Project | Description |
|---|---|
| CardKS · Contributor | Residual key-structure modeling for long-horizon decisions across GuanDan, DouDizhu, and Gin Rummy. |
| DanKS · Contributor | PPO self-play GuanDan agent with structure-aware retrieval and candidate-conditioned actor-critic learning. |
| Paper Figure Skill | Reconstructs paper figures as editable PowerPoint diagrams with reusable assets and visual verification. |
| SMRL | PyTorch/DGL implementation of spatial meta-learning for unseen geographic entities. |
| STMetaT | Spatio-temporal meta-learning for multi-view trajectory representations. |

