Haofu Liao

Principal Applied Scientist,
AWS Agentic AI

liaohaofu [at] gmail [dot] com

Haofu Liao 廖昊夫

I am a tech lead and Principal Applied Scientist at AWS Agentic AI, and previously led document-understanding research at AWS AI Labs. I received my Ph.D. in Computer Science from the University of Rochester, where I worked with Jiebo Luo, my M.S. in Electrical and Computer Engineering from Northeastern University in Boston, and my B.Eng. from the Beijing University of Posts and Telecommunications.

My work centers on agentic AI: building autonomous agents and the foundation models behind them that can interpret multimodal inputs, reason over complex context, and act to complete real-world tasks. My research spans agentic systems, large action models, and the orchestration that turns models into reliable agents.

At AWS, I build the agentic systems and train the large action models behind Amazon Quick, where I lead the science for its workflow automation (Quick Automate and Quick Flows). Before that, I led the science behind AWS products including Bedrock Data Automation and Textract, and developed document-understanding methods such as DocTr and DocKD. My earlier research in medical image computing introduced methods like DuDoNet and ADN for CT artifact reduction. Altogether, my work has been cited over 2,600 times with an h-index of 26.

I typically host one summer intern each year, so if you are interested in an internship with AWS Agentic AI, feel free to reach out.

News

Selected Research

See all research →
Amazon Quick

Amazon Quick Workflow Automation

2024 – Present

I lead the science for agentic workflow automation in Amazon Quick, AWS's agentic AI platform for workplace automation. I started with GUI automation, building a vision-language large action model from scratch in 2024 when no comparable solution existed in the industry, and directed teams across data, mid-training, and post-training to take it from research to production. I have since expanded to lead the science across the full stack of workflow automation, where agents complete tasks over user interfaces, APIs, tools, and code, spanning Quick Automate for complex multi-step processes and Quick Flows for no-code automation of routine tasks. This work is reflected in our research on GUI grounding (CVPR 2026) and efficient web automation (Findings of ACL 2025).

AWS DevOps Agent

AWS DevOps Agent Release Management

2025 – Present

I am the tech lead for autonomous release testing in AWS DevOps Agent's release management capabilities. The agent generates and runs change-specific tests for web and API applications in production-like environments, helping teams assess functional correctness, behavioral regressions, and integration risks before changes reach production.

Amazon Bedrock Data Automation

Amazon Bedrock Data Automation

2023 – 2024

I was a tech lead for Amazon Bedrock Data Automation, the AWS service that turns unstructured multimodal content (documents, images, video, and audio) into structured insights for generative-AI applications, where I led model training for several core components and helped bring the service to public preview. On the problem of making compact document-understanding models generalize to unseen formats, I co-led DocKD (co-first author, EMNLP 2024), a knowledge-distillation method that feeds large language models structured document elements (key-value pairs, layouts, and descriptions) to generate higher-quality synthetic training data. Models trained only on DocKD data match human-annotated data in-domain and surpass it out-of-domain.

Amazon Textract

Amazon Textract Analyze Expense

2021 – 2023

I was a tech lead for Amazon Textract Analyze Expense, the capability that extracts fields and line items from invoices and receipts without templates. On this structured-extraction problem I led DocTr (first author, ICCV 2023), a document transformer that recasts information extraction as anchor-based detection: each entity is represented by an anchor word and a bounding box, with relations captured through anchor-word associations, replacing the fragile token tagging and brittle graph decoding of prior methods. DocTr achieved state-of-the-art results across three extraction benchmarks.

Selected Publications

See all publications →
GUI grounding figure

Towards GUI Agents: Vision-Language Diffusion Models for GUI Grounding

Shrinidhi Kumbhar, Haofu Liao, Srikar Appalaraju, Kunwar Yashraj Singh.

IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2026.

Web history compressor architecture

Turbocharging Web Automation: The Impact of Compressed History States

Xiyue Zhu, Peng Tang, Haofu Liao, Srikar Appalaraju.

Findings of the Association for Computational Linguistics (ACL), 2025.

DocKD figure

DocKD: Knowledge Distillation from LLMs for Open-World Document Understanding Models

Sungnyun Kim*, Haofu Liao*, Srikar Appalaraju, Peng Tang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan, Stefano Soatto.

Conference on Empirical Methods in Natural Language Processing (EMNLP), 2024.

DocTr figure

DocTr: Document Transformer for Structured Information Extraction in Documents

Haofu Liao, Aruni RoyChowdhury, Weijian Li, Ankan Bansal, Yuting Zhang, Zhuowen Tu, Ravi Kumar Satzoda, R. Manmatha, Vijay Mahadevan.

IEEE/CVF International Conference on Computer Vision (ICCV), 2023.

Deep Network Design for Medical Image Computing book cover

Deep Network Design for Medical Image Computing: Principles and Applications

Haofu Liao, S. Kevin Zhou, Jiebo Luo.

Academic Press, 2022.  Book