Note
AI-generated profile summary. This overview was synthesized from recurring themes across my recent and long-running public and private repositories. Private repository names and implementation details are intentionally omitted.
My work follows a broad path from visual learning research toward multimodal perception and embodied systems:
- Visual intelligence — image segmentation, object detection, visual grounding, open-vocabulary perception, and VLM evaluation.
- Data-centric AI — synthetic data generation, automatic annotation, quality control, augmentation, and dataset selection.
- Embodied AI — indoor robotics, RGB-D perception, ROS 2, navigation, and simulation in Isaac Sim.
- 3D scene infrastructure — USD pipelines, scene conversion, collision geometry, and simulation-ready environments.
The longer arc began with biomedical and industrial vision—semi-supervised and continual segmentation, dense prediction, depth, and autofocus—and has gradually moved closer to systems that perceive, reason, and act.
Outside the main research thread, I tend to make tools around whatever gets in the way:
- Developer tools and editor extensions
- Small desktop apps, CLIs, and automation
- Generative image and video workflows
- Bots, infrastructure, Docker, and network utilities
Selected public projects:
- LabelEditor for VS Code — image annotation with LabelMe and SAM-assisted masks.
- MiniMax H3 Bot — a multi-GPU image-to-video service with ComfyUI and Telegram integration.
- Sidebrowser — a compact side-panel browser for Windows.
- Neon Postgres Sync — two-way text synchronization inside VS Code.

