Add superseded/exploratory BEFT and HiRA results
Browse files
imagegen/beft--flux2-klein-default-original.json
ADDED
|
@@ -0,0 +1,18 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"peft_method": "beft",
|
| 3 |
+
"base_model": "flux2-klein",
|
| 4 |
+
"variant": "default-pre-to_out-fix",
|
| 5 |
+
"target_modules": ["to_q", "to_k", "to_v", "to_out.0", "add_q_proj", "add_k_proj", "add_v_proj", "to_add_out", "to_qkv_mlp_proj", "linear_in", "linear_out"],
|
| 6 |
+
"batch_size": 2,
|
| 7 |
+
"lr": 1e-4,
|
| 8 |
+
"test_dino_similarity": 0.5227,
|
| 9 |
+
"test_drift": 0.0732,
|
| 10 |
+
"comparison_baseline": {
|
| 11 |
+
"config": "beft/flux2-klein-default (with single-stream to_out added, adopted)",
|
| 12 |
+
"test_dino_similarity": 0.6015,
|
| 13 |
+
"test_drift": 0.1082
|
| 14 |
+
},
|
| 15 |
+
"conclusion": "Superseded by a target_modules fix: the single-stream transformer blocks' `to_out` is a bare Linear (unlike the double-stream blocks' `to_out.0` inside a ModuleList), so it was silently uncovered here. Adding it raised DINOv2 similarity by +15% for +61,440 trainable params at negligible runtime cost. Kept as the pre-fix reference point.",
|
| 16 |
+
"note": "BEFT has no rank/capacity hyperparameter to tune, so target_modules coverage was the only lever available to address the underfitting the maintainer flagged. See PEFT discussion #3522 and PR #3663.",
|
| 17 |
+
"environment": "Single RTX 5070 Ti, 16GB VRAM"
|
| 18 |
+
}
|
imagegen/hira--flux2-klein-default-lr1e-4.json
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"peft_method": "hira",
|
| 3 |
+
"base_model": "flux2-klein",
|
| 4 |
+
"variant": "lr1e-4",
|
| 5 |
+
"rank": 32,
|
| 6 |
+
"hira_dropout": 0.0,
|
| 7 |
+
"batch_size": 2,
|
| 8 |
+
"lr": 1e-4,
|
| 9 |
+
"test_dino_similarity": 0.5550,
|
| 10 |
+
"test_drift": 0.1454,
|
| 11 |
+
"comparison_baseline": {
|
| 12 |
+
"config": "hira/flux2-klein-default (lr=5e-4, adopted)",
|
| 13 |
+
"test_dino_similarity": 0.7075,
|
| 14 |
+
"test_drift": 0.2626
|
| 15 |
+
},
|
| 16 |
+
"conclusion": "The shared default LR (1e-4) undertrains HiRA on this task. Tuning the LR up to 5e-4 closed the gap with LoRA and slightly exceeded it, at near-identical drift. Kept as the pre-tuning reference point showing the shared-default result.",
|
| 17 |
+
"note": "Validation DINOv2 curve looked unconverged at step 750 (climbed unevenly, jumped from 0.418 at step 600 to 0.583 at step 700) -- same shape recurred at lr=1.53e-4, suggesting the issue is convergence speed at this step budget rather than LR alone. A transient CUDA OOM warning appeared early in training but auto-recovered and did not affect the run. See PEFT discussion #3522 and PR #3661.",
|
| 18 |
+
"environment": "Single RTX 5070 Ti, 16GB VRAM"
|
| 19 |
+
}
|
imagegen/hira--flux2-klein-lr-tuned-1.53e-4.json
ADDED
|
@@ -0,0 +1,19 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"peft_method": "hira",
|
| 3 |
+
"base_model": "flux2-klein",
|
| 4 |
+
"variant": "lr1.53e-4",
|
| 5 |
+
"rank": 32,
|
| 6 |
+
"hira_dropout": 0.0,
|
| 7 |
+
"batch_size": 2,
|
| 8 |
+
"lr": 1.53e-4,
|
| 9 |
+
"test_dino_similarity": 0.6333,
|
| 10 |
+
"test_drift": 0.2285,
|
| 11 |
+
"comparison_baseline": {
|
| 12 |
+
"config": "hira/flux2-klein-default (lr=5e-4, adopted)",
|
| 13 |
+
"test_dino_similarity": 0.7075,
|
| 14 |
+
"test_drift": 0.2626
|
| 15 |
+
},
|
| 16 |
+
"conclusion": "Best value from a 20-trial Optuna sweep within the HiRA paper's suggested [1e-4, 2e-4] range. Confirmed the direction (higher LR helps) but a further push to 5e-4, beyond the paper's suggested range, improved results more. Not added as a separate benchmark entry per maintainer guidance (avoid too many near-duplicate LR points); kept here as the intermediate step in the tuning trajectory.",
|
| 17 |
+
"note": "Same unconverged-looking validation curve shape as the lr=1e-4 run (plateau then late jump), suggesting convergence speed at the shared 750-step budget is a factor independent of the specific LR tried. See PEFT discussion #3522 and PR #3661.",
|
| 18 |
+
"environment": "Single RTX 5070 Ti, 16GB VRAM"
|
| 19 |
+
}
|