Skip to content
View Thanx01's full-sized avatar

Block or report Thanx01

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Thanx01/README.md
Zhouzheng Xu — LLM Post-Training, Web Agents, and Game AI

Email GitHub

Research interests center on LLM post-training and sequential decision-making, with current work on WebAgent systems and card-game AI. Core methods include supervised fine-tuning, reinforcement learning, self-play, and Monte Carlo tree search in imperfect-information games; earlier work focused on diffusion-based generative modeling and distributed training.

Experience

Kingsoft AI Research · Kingsoft AI
2025.07 — Present
WebAgent systems and card-game AI; supervised fine-tuning, reinforcement learning, self-play, Monte Carlo tree search, and imperfect-information games.
Migu Music AIGC Research · Migu Music
2025.04 — 2025.07
Diffusion-based generative modeling and large-scale training with DeepSpeed, supervised fine-tuning, and reinforcement learning.

Research & Engineering

LLM Post-Training Supervised fine-tuning Reinforcement learning DeepSpeed
Agents WebAgent Web navigation
Game AI Self-play Monte Carlo tree search Imperfect-information games
Generative Modeling Diffusion models AIGC
Research Code Python PyTorch DGL

Publications

Publication Venue Research topic
SMRL · Paper IEEE Transactions on Neural Networks and Learning Systems, 2026 Spatial meta-learning for unseen geographic entities
STMetaT · Paper Knowledge-Based Systems, 2025 Spatio-temporal meta-learning for trajectory representation

Open Source

Project Description
CardKS · Contributor Residual key-structure modeling for long-horizon decisions across GuanDan, DouDizhu, and Gin Rummy.
DanKS · Contributor PPO self-play GuanDan agent with structure-aware retrieval and candidate-conditioned actor-critic learning.
Paper Figure Skill Reconstructs paper figures as editable PowerPoint diagrams with reusable assets and visual verification.
SMRL PyTorch/DGL implementation of spatial meta-learning for unseen geographic entities.
STMetaT Spatio-temporal meta-learning for multi-view trajectory representations.

GitHub Activity

Thanx01 contribution activity Thanx01 GitHub statistics Thanx01 languages by commit

Pinned Loading

  1. Calix-L/CardKS Calix-L/CardKS Public

    Paper and ecosystem hub for residual key-structure modeling across GuanDan, DouDizhu, and Gin Rummy.

    Python 2 1

  2. paper-figure-skill paper-figure-skill Public

    Agent Skill for turning papers or reference figures into editable PowerPoint diagrams with reusable assets and visual verification.

    JavaScript 2

  3. SMRL SMRL Public

    Official PyTorch/DGL implementation of SMRL (IEEE TNNLS 2026): spatial meta-learning for unseen geographic entities.

    Python 1

  4. STMetaT STMetaT Public

    Official implementation of STMetaT (Knowledge-Based Systems 2025): spatio-temporal meta-learning for trajectory representations.

    Python 1

  5. Calix-L/DanKS Calix-L/DanKS Public

    RL‑Empowered Small‑Scale Competitive Guandan Agent

    Python 428 11