Sergasgr commited on
Commit
a96120c
·
verified ·
1 Parent(s): 982d66e

Add superseded/exploratory BEFT and HiRA results

Browse files
imagegen/beft--flux2-klein-default-original.json ADDED
@@ -0,0 +1,18 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "peft_method": "beft",
3
+ "base_model": "flux2-klein",
4
+ "variant": "default-pre-to_out-fix",
5
+ "target_modules": ["to_q", "to_k", "to_v", "to_out.0", "add_q_proj", "add_k_proj", "add_v_proj", "to_add_out", "to_qkv_mlp_proj", "linear_in", "linear_out"],
6
+ "batch_size": 2,
7
+ "lr": 1e-4,
8
+ "test_dino_similarity": 0.5227,
9
+ "test_drift": 0.0732,
10
+ "comparison_baseline": {
11
+ "config": "beft/flux2-klein-default (with single-stream to_out added, adopted)",
12
+ "test_dino_similarity": 0.6015,
13
+ "test_drift": 0.1082
14
+ },
15
+ "conclusion": "Superseded by a target_modules fix: the single-stream transformer blocks' `to_out` is a bare Linear (unlike the double-stream blocks' `to_out.0` inside a ModuleList), so it was silently uncovered here. Adding it raised DINOv2 similarity by +15% for +61,440 trainable params at negligible runtime cost. Kept as the pre-fix reference point.",
16
+ "note": "BEFT has no rank/capacity hyperparameter to tune, so target_modules coverage was the only lever available to address the underfitting the maintainer flagged. See PEFT discussion #3522 and PR #3663.",
17
+ "environment": "Single RTX 5070 Ti, 16GB VRAM"
18
+ }
imagegen/hira--flux2-klein-default-lr1e-4.json ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "peft_method": "hira",
3
+ "base_model": "flux2-klein",
4
+ "variant": "lr1e-4",
5
+ "rank": 32,
6
+ "hira_dropout": 0.0,
7
+ "batch_size": 2,
8
+ "lr": 1e-4,
9
+ "test_dino_similarity": 0.5550,
10
+ "test_drift": 0.1454,
11
+ "comparison_baseline": {
12
+ "config": "hira/flux2-klein-default (lr=5e-4, adopted)",
13
+ "test_dino_similarity": 0.7075,
14
+ "test_drift": 0.2626
15
+ },
16
+ "conclusion": "The shared default LR (1e-4) undertrains HiRA on this task. Tuning the LR up to 5e-4 closed the gap with LoRA and slightly exceeded it, at near-identical drift. Kept as the pre-tuning reference point showing the shared-default result.",
17
+ "note": "Validation DINOv2 curve looked unconverged at step 750 (climbed unevenly, jumped from 0.418 at step 600 to 0.583 at step 700) -- same shape recurred at lr=1.53e-4, suggesting the issue is convergence speed at this step budget rather than LR alone. A transient CUDA OOM warning appeared early in training but auto-recovered and did not affect the run. See PEFT discussion #3522 and PR #3661.",
18
+ "environment": "Single RTX 5070 Ti, 16GB VRAM"
19
+ }
imagegen/hira--flux2-klein-lr-tuned-1.53e-4.json ADDED
@@ -0,0 +1,19 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "peft_method": "hira",
3
+ "base_model": "flux2-klein",
4
+ "variant": "lr1.53e-4",
5
+ "rank": 32,
6
+ "hira_dropout": 0.0,
7
+ "batch_size": 2,
8
+ "lr": 1.53e-4,
9
+ "test_dino_similarity": 0.6333,
10
+ "test_drift": 0.2285,
11
+ "comparison_baseline": {
12
+ "config": "hira/flux2-klein-default (lr=5e-4, adopted)",
13
+ "test_dino_similarity": 0.7075,
14
+ "test_drift": 0.2626
15
+ },
16
+ "conclusion": "Best value from a 20-trial Optuna sweep within the HiRA paper's suggested [1e-4, 2e-4] range. Confirmed the direction (higher LR helps) but a further push to 5e-4, beyond the paper's suggested range, improved results more. Not added as a separate benchmark entry per maintainer guidance (avoid too many near-duplicate LR points); kept here as the intermediate step in the tuning trajectory.",
17
+ "note": "Same unconverged-looking validation curve shape as the lr=1e-4 run (plateau then late jump), suggesting convergence speed at the shared 750-step budget is a factor independent of the specific LR tried. See PEFT discussion #3522 and PR #3661.",
18
+ "environment": "Single RTX 5070 Ti, 16GB VRAM"
19
+ }