I am an AI algorithm engineer at Malanshan Audio-Video Laboratory. I train agent foundations for audio and video: private-domain knowledge, affective dialogue, and the pipelines that keep generated media consistent. 我是马栏山音视频实验室的 AI 算法工程师。目前在做音视频智能体基座,包括私域知识、情感对话,以及让生成内容保持一致的数据管线。

I studied software engineering and computer technology at Beijing Institute of Technology. Before this role I led a spoken-misinformation study with CUHK-Shenzhen and Huawei. 本科与硕士都在北京理工大学,分别是软件工程和计算机技术。在此之前,我在香港中文大学(深圳)与华为的合作研究中负责语音虚假信息检测。

10Patent filings独立申请专利
2Papers已发表论文
10k+Stars on projects I joined参与项目累计 Star
85.7GPA, top 5% of majorGPA,专业前 5%

Selected work工作

Malanshan Audio-Video Laboratory · AI Algorithm Engineer · 2025.06 — Present. Ten patent applications, filed independently. 马栏山音视频实验室 · AI 算法工程师 · 2025.06 — 至今。已独立申请专利 10 篇。

01 4 filings专利 4 篇

Intelligent assistant foundation智能助手 AI-Agent 基座

Trained a foundation LLM and tightened intent recognition and procedure modules with RLHF and DPO. Adaptive workflows and VLA-style models power a classified, private-domain RAG assistant. A daily insight agent cleans external knowledge, scores relevance, and drafts the briefing deck plus a live view. For the China Film Science Research Institute text-to-film project, ReAct splits scripts so generation stays efficient and character control stays consistent. 训练基础 LLM,用 RLHF 和 DPO 加强意图识别与规程模块。自适应工作流和 VLA 等模型支撑内部分类的私域 RAG。洞察 Agent 按天清洗外部知识、计算关联度,并生成简报大纲和动态可视化。在中国电影科学研究所的文生电影项目里,用 ReAct 拆分剧本,提高生成效率,并保持人物规控一致。

02 5 filings专利 5 篇

Realtime affective agent实时多模态情感交互

A VLM drives proactive multimodal interaction. Grammar checking cut scheduled-task latency, and the loop runs on low-compute edge devices. Face and emotion detectors, tuned together with the VLM, let the dialogue inject emotion and control its intensity. 用 VLM 驱动主动的多模态交互。文法检测降低了定时任务时延,交互循环可以跑在小算力端侧设备上。人脸与情绪检测和 VLM 一起优化后,对话可以按情感注入并调节强度。

03 1 filing专利 1 篇

Data-flow foundation数据流转基座

Following a data-generation paradigm from Carnegie Mellon University, I trained a synthesis model for natural dialogue and classification data, then used RL and DPO to improve naturalness and usability. An agent pipeline self-corrects real-scene ASR annotation and cut about ¥800,000 in manual correction cost on the related project. 参考卡内基梅隆大学的数据生成范式,自训练了面向自然对话和分类数据的合成模型,再用 RL 与 DPO 提高自然度和可用性。另一条 AI-Agent 管线会自主纠正真实场景的 ASR 标注,相关项目大约减少了 80 万人工纠错成本。

Education教育

Beijing Institute of Technology · Project 985 · 211 · Double First-Class · Haidian, Beijing 北京理工大学 · 第一批直属 985 · 211 · 双一流 · 北京海淀

2022.09 — 2025.06

M.Eng. Computer Technology计算机技术 · 硕士

GPA 85.7. Top 5% of the major. Minority Scholarship and Excellent Scholarship.GPA 85.7,专业前 5%。少数民族奖学金、优秀奖学金。

2018.09 — 2022.06

B.Eng. Software Engineering软件工程 · 本科

Top 5% of the major across both degrees.本硕专业排名均位于前 5%。

CET-4 · CET-6 448 · GRE V 135 + Q 166

Python · C++ · Java · Rust

毕业照
Graduation毕业

Research科研与实习

2024.07 — 2024.11

CUHK-Shenzhen × Huawei Singapore香港中文大学(深圳)× 华为(新加坡)

Research lead · Spoken misinformation研究组长 · 语音虚假信息检测

Proposed the task, built the dataset with TTS, and assembled a three-stage detection pipeline. I cleaned the text and speech, generated the audio, and wrote the paper. 提出该任务,用 TTS 建成数据集,并搭建三阶段检测流程。负责文本与语音清洗、音频生成和论文撰写。

2024.03 — 2024.07

Tongji Psychology × Toyota × FORVIA同济心理 × 丰田 × FORVIA

Algorithm engineer · Personality model and text-to-video算法工程师 · 人格大模型与文生视频

Fine-tuned Qwen, built the data pipeline, and tuned classification and dialogue models with automated prompts. I trained the psychological-mode and intervention-mode classifiers, and the personality model. 微调 Qwen,完成数据管线,以及分类模型和对话模型。负责心理模式、干预模式分类,和人格大模型训练。

2023.02 — 2023.11

NIO蔚来汽车

Algorithm engineer · Test-case generation算法工程师 · 测试用例生成

Turned requirement documents into test cases with a large model, RAG, and a database. I fine-tuned LLaMA, designed the prompts and chain-of-thought, and built the retrieval stack and frontend APIs. 用大模型、RAG 和数据库,把需求文档生成测试用例。负责 LLaMA 微调、多级 prompt、思维链、检索系统,以及前端接口。

Papers论文与荣誉

IEEE SLT 2024 · Best Student Paper · Equal contribution最佳学生论文 · 共同一作

SpMis: An Investigation of Synthetic Spoken Misinformation Detection

Computer Science · CCF T2 · Student first author《计算机科学》· CCF T2 · 学生一作

A Single-Stage Unsupervised Visible-Infrared Person Re-Identification Method一种单阶段无监督可见光-红外跨模态行人重识别方法

Open source开源

Contributions to OpenMMLab and Amphion. Those projects have accumulated more than 10k stars. 参与 OpenMMLab 与 Amphion,相关项目累计超过 10k stars。

  • VocalGuard — awarded at the WeBank FinTech Competition, Shenzhen.VocalGuard — 深圳微众银行金融科技大赛获奖。
  • Minority Scholarship and Excellent Scholarship, BIT.北京理工大学少数民族奖学金、优秀奖学金。
  • Bronze Award, 15th Challenge Cup, 2020.第十五届“挑战杯”铜奖,2020。
  • Second Prize, Xindao Cup Yunnan sand-table competition, 2020.“新道杯”云南省沙盘比赛二等奖,2020。

Life生活

From Lijiang. Most of these were taken between the lab and the mountains. 籍贯丽江。这些照片大多拍在实验室和山之间。

山脊上的和任强
雪场
高原
草地
山谷
夜里