UIPress / START_HERE.md
DesonDai's picture
Add files using upload-large-folder tool
bac741f verified
|
Raw History Blame Contribute Delete
11.2 kB

UIPress: 快速开始指南

硬件: 6 × NVIDIA A40 (48GB) 统一模型: Qwen3-VL-8B-Instruct 预计总时间: 2-3 天

当前任务状态(已按实际运行标记)

  • 1. 环境配置:依赖与基础环境已可用。
  • 2. 数据准备:训练与评估数据已就位并可被脚本读取。
  • 3.1 Smoke Test:单卡冒烟已跑通。
  • 3.2 正式训练:已完成(最新日志:Epoch 4: avg_loss=0.2288,并保存 checkpoints/optical/epoch4.pt)。
  • 3.3 选择最佳 checkpoint:待最终确认(当前已有 epoch0/epoch1/epoch4/latest/best,建议按最低 loss 更新 best.pt)。
  • 4.1 一键并行评估(50样本):已完成。
  • 4.2 逐个评估(已完成项):baseline、visionzip-256、visionzip-128、efficientui(prune=0.6)、uipress-256。
  • 4.2 逐个评估(补齐项):resolution(230400/1003520)、visionzip-64、efficientui(prune=0.8) 已完成。
  • 4.3 计算 CLIP 分数:已完成(results/benchmark/all_clip_scores.json)。
  • SSIM/bootstrap:已完成(results/benchmark/ssim_scores.json、results/benchmark/bootstrap_ci.json)。
  • element_analysis:已完成(results/element_analysis.json)。
  • case study:尚未执行。

执行约定(当前)

  • 所有新任务统一使用 nohup 后台启动,并写入 logs/*.nohup.log。

现在可继续做(按优先级)

  1. 最终 checkpoint 选择(推荐先做)
for f in checkpoints/optical/epoch*.pt; do
  echo -n "$f: "
  python -c "import torch; c=torch.load('$f', map_location='cpu'); print(f'loss={c[\"loss\"]:.4f}')"
done
# 将最佳 epoch 覆盖为 best.pt
# cp checkpoints/optical/epochX.pt checkpoints/optical/best.pt
  1. 补跑 case study(nohup)
nohup bash -lc 'PYTHONPATH=. python scripts/step_case_study.py' \
  > logs/case_study.nohup.log 2>&1 < /dev/null &
  1. 可选:重跑失败样本渲染后再算 SSIM(例如 129/130)
nohup bash -lc 'PYTHONPATH=. python scripts/step_ssim_bootstrap.py --benchmark_dir results/benchmark --ref_dir data/ref_screenshots' \
  > logs/ssim_rerun.nohup.log 2>&1 < /dev/null &

0. 项目结构

UIPress/
├── models/
│   ├── __init__.py
│   └── optical_compressor.py      # UIPress 光学压缩模块
├── scripts/
│   ├── train_compressor.py        # 训练脚本 (DDP)
│   ├── eval_all.py                # 统一评估脚本 (4种方法)
│   ├── run_all_evals.sh           # 一键并行评估
│   ├── step_clip_batch.py         # CLIP 评分 (已有)
│   ├── step_ssim_bootstrap.py     # SSIM + CI (已有)
│   ├── step_element_analysis.py   # HTML 结构分析 (已有)
│   ├── step_case_study.py         # 可视化 (已有)
│   ├── download_data.py           # 数据下载
│   └── download_websight.py       # WebSight 下载
├── data/
│   ├── design2code/               # 测试集
│   ├── ref_screenshots/           # 参考截图
│   └── websight/                  # 训练数据
├── results/
│   └── comparison/                # 新实验结果输出
├── checkpoints/                   # 训练 checkpoint
├── logs/                          # 日志
└── COLM2026/                      # 论文

1. 环境配置

# 创建 conda 环境
conda create -n uipress python=3.11 -y
conda activate uipress

# 安装依赖
pip install -r requirements.txt

# 安装 Playwright (HTML 渲染用)
playwright install chromium

# 验证
python -c "from transformers import Qwen3VLForConditionalGeneration; print('OK')"

2. 数据准备

2.1 训练数据 (WebSight)

# 下载 WebSight 子集 (50K 样本用于训练)
PYTHONPATH=. python scripts/download_websight.py --max_samples 50000

# 或者如果已有 websight 数据:
ls data/websight/  # 确认存在

2.2 测试数据 (Design2Code)

# 确认 485 张参考截图已就位
ls data/ref_screenshots/ | wc -l
# 应该显示 485

3. 训练 UIPress 光学压缩器

3.1 Smoke Test (单卡, 5分钟)

先验证代码能跑通:

CUDA_VISIBLE_DEVICES=0 python scripts/train_compressor.py \
    --max_samples 20 \
    --epochs 1 \
    --batch_size 1 \
    --grad_accum 1 \
    --target_tokens 256

预期输出:

  • Compressor params: ~30M
  • LoRA params: ~6M
  • loss 应该在 3-5 范围,不是 NaN
  • 显存 ~25-30GB

3.2 正式训练 (6卡 DDP, 8-16小时)

mkdir -p logs checkpoints/optical

torchrun --nproc_per_node=6 scripts/train_compressor.py \
    --data_dir data/websight \
    --max_samples 50000 \
    --epochs 5 \
    --batch_size 1 \
    --grad_accum 8 \
    --lr_compressor 2e-4 \
    --lr_lora 2e-5 \
    --target_tokens 256 \
    --output_dir checkpoints/optical \
    2>&1 | tee logs/train_compressor.log

训练参数说明:

参数 值 说明
effective batch 1×8×6=48 batch_size × grad_accum × GPUs
lr_compressor 2e-4 压缩模块学习率 (新模块, 较大)
lr_lora 2e-5 LLM LoRA 学习率 (微调, 较小)
target_tokens 256 压缩到 256 tokens (16×16 grid)
显存/卡 ~26GB A40 48GB 足够

3.3 选择最佳 checkpoint

训练完成后, 挑 loss 最低的 epoch:

# 查看每个 epoch 的 loss
for f in checkpoints/optical/epoch*.pt; do
    echo -n "$f: "
    python -c "import torch; c=torch.load('$f',map_location='cpu'); print(f'loss={c[\"loss\"]:.4f}')"
done

# 软链接最佳
cp checkpoints/optical/epoch3.pt checkpoints/optical/best.pt  # 替换为实际最佳

4. 评估 (所有方法)

4.1 一键并行评估 (6卡, 3-4小时)

mkdir -p logs

# 50 样本快速验证
chmod +x scripts/run_all_evals.sh
./scripts/run_all_evals.sh 50

# 全量 485 样本
./scripts/run_all_evals.sh 485

4.2 逐个评估 (调试用)

# 方法 1: Baseline (无压缩)
CUDA_VISIBLE_DEVICES=0 python scripts/eval_all.py \
    --method baseline --max_samples 50

# 方法 2: 分辨率缩放 (多个档位)
CUDA_VISIBLE_DEVICES=1 python scripts/eval_all.py \
    --method resolution --max_pixels 230400 --max_samples 50   # ~256px → ~720 tokens

CUDA_VISIBLE_DEVICES=1 python scripts/eval_all.py \
    --method resolution --max_pixels 1003520 --max_samples 50  # ~1k → ~3000 tokens

# 方法 3: VisionZip (多个 token 数)
CUDA_VISIBLE_DEVICES=2 python scripts/eval_all.py \
    --method visionzip --keep_tokens 256 --max_samples 50

CUDA_VISIBLE_DEVICES=2 python scripts/eval_all.py \
    --method visionzip --keep_tokens 128 --max_samples 50

CUDA_VISIBLE_DEVICES=2 python scripts/eval_all.py \
    --method visionzip --keep_tokens 64 --max_samples 50

# 方法 4: EfficientUICoder 策略 (多个剪枝率)
CUDA_VISIBLE_DEVICES=3 python scripts/eval_all.py \
    --method efficientui --prune_ratio 0.6 --max_samples 50

CUDA_VISIBLE_DEVICES=3 python scripts/eval_all.py \
    --method efficientui --prune_ratio 0.8 --max_samples 50

# 方法 5: UIPress 光学压缩 (多个 token 数)
CUDA_VISIBLE_DEVICES=4 python scripts/eval_all.py \
    --method uipress --checkpoint checkpoints/optical/best.pt \
    --target_tokens 256 --max_samples 50

4.3 计算 CLIP 分数

评估脚本只生成 HTML, 还需渲染+计算 CLIP:

# 对所有结果批量计算 CLIP
CUDA_VISIBLE_DEVICES=0 PYTHONPATH=. python scripts/step_clip_batch.py

5. 结果汇总

评估完成后, 结果在 results/comparison/ 下:

results/comparison/
├── qwen3_full/
│   ├── html_predictions/     # 生成的 HTML 文件
│   ├── summary.json          # 汇总 (tokens, 延迟, 显存)
│   └── per_sample.json       # 逐样本结果
├── visionzip_256/
├── visionzip_128/
├── efficientui_prune60/
├── uipress_256/
└── qwen3_res_230400/

当前实测结果(50样本,已完成)

方法 n_success 平均视觉 tokens 平均延迟 平均峰值显存
Qwen3-VL full 50/50 7299.2 90.83s 16.96GB
VisionZip-256 50/50 256.0 110.44s 17.07GB
VisionZip-128 50/50 128.0 108.41s 17.07GB
VisionZip-64 50/50 64.0 92.04s 17.04GB
EfficientUI-60% 50/50 729.9 99.52s 17.02GB
EfficientUI-80% 50/50 364.0 102.37s 17.05GB
Qwen3-Res-230400 50/50 844.5 94.00s 16.74GB
Qwen3-Res-1003520 50/50 3747.5 76.96s 16.79GB
UIPress-256 50/50 256.0 52.52s 17.31GB

当前 CLIP / SSIM(50样本)

方法 CLIP SSIM
qwen3_res_230400 0.7768 0.6592
qwen3_res_1003520 0.7750 0.6612
qwen3_full 0.7563 0.6647*
efficientui_prune60 0.7523 0.6487
efficientui_prune80 0.7380 0.6232
visionzip_256 0.7333 0.6489
visionzip_128 0.7245 0.6461
uipress_256 0.7232 0.6323
visionzip_64 0.7197 0.6452

* qwen3_full 与 qwen3_res_1003520 在 SSIM 渲染阶段各有少量超时样本(统计时已按可用样本数计算)。

当前还能马上做什么(全部可立即启动)

  1. 恢复训练(优先):当前训练在 E2 S4272/10000 中断,可从 checkpoints/optical/latest.pt 继续。
  2. 补齐剩余评估档位:resolution、visionzip-64、efficientui(prune=0.8)。
  3. 后处理与统计:运行 step_clip_batch.py、step_ssim_bootstrap.py,然后做 step_element_analysis.py / step_case_study.py。

预期对比表

方法 Tokens CLIP (预期) 延迟 显存
Qwen3-VL full ~6700 0.776 ~80s 17GB
分辨率缩放 256px ~720 0.765 ~90s 17GB
VisionZip-256 256 ~0.76-0.78 ~30s ~12GB
VisionZip-128 128 ~0.74-0.76 ~25s ~10GB
EfficientUI-60% ~2700 ~0.76-0.78 ~60s 17GB
UIPress-256 256 目标>0.78 ~30s ~12GB

6. 故障排除

OOM

# 检查 GPU 显存
nvidia-smi --query-gpu=index,memory.used,memory.free --format=csv

# 训练 OOM → 减小 batch_size 或增大 grad_accum
# 评估 OOM → 减小 max_new_tokens (默认 4096)

训练 loss 不降

  1. 检查数据: python -c "from datasets import load_from_disk; d=load_from_disk('data/websight'); print(len(d))"
  2. 降低 lr_compressor 到 1e-4
  3. 检查 WebSight 数据是否有 code 或 text 字段

CLIP 评分为 0 或 NaN

  1. 检查 HTML 是否为空: head results/comparison/*/html_predictions/0.html
  2. 检查 Playwright 是否安装: playwright install chromium
  3. 检查参考截图: ls data/ref_screenshots/ | head

7. 时间线

天 任务 GPU
Day 1 上午 环境配置 + Smoke Test 1 GPU
Day 1 下午-晚 正式训练 (8-16h) 6 GPUs DDP
Day 2 上午 选 checkpoint + 并行评估 6 GPUs 各跑一个方法
Day 2 下午 CLIP 批量评分 + SSIM 1 GPU
Day 3 结果分析 + 补充实验 按需