Christian Medina
|
3892ea7fec
|
feat: test with BnB 4-bit quantization to force proper loading
|
2026-07-02 12:23:34 -04:00 |
|
Christian Medina
|
c5a3b87eed
|
feat: test both strategies - device_map=auto and FSDP
|
2026-07-02 12:19:05 -04:00 |
|
Christian Medina
|
7f498134f2
|
fix: FSDP loads model to GPU first, then shards across GPUs
|
2026-07-02 12:17:44 -04:00 |
|
Christian Medina
|
f9c748706f
|
fix: use already-quantized 4-bit model for training (no BnB needed)
|
2026-07-02 12:09:59 -04:00 |
|
Christian Medina
|
3e26d7d37e
|
fix: use already-quantized 4-bit model for test
|
2026-07-02 12:09:49 -04:00 |
|
Christian Medina
|
e3ea60e6c6
|
fix: clarify distribution pattern detection in test script
|
2026-07-02 12:04:54 -04:00 |
|
Christian Medina
|
3049aa9b0a
|
feat: add test script to verify model loading and GPU distribution
|
2026-07-02 12:04:04 -04:00 |
|
Christian Medina
|
14ef1a07c4
|
fix: remove sync_module_states=True (model is on CPU)
|
2026-07-02 11:55:21 -04:00 |
|
Christian Medina
|
5a13ca1d1c
|
feat: FSDP training failures now fallback to next strategy
|
2026-07-02 11:48:03 -04:00 |
|
Christian Medina
|
ac1417567c
|
fix: fix another indentation error
|
2026-07-02 11:38:54 -04:00 |
|
Christian Medina
|
243223d899
|
fix: fix indentation error in strategy 3
|
2026-07-02 11:38:37 -04:00 |
|
Christian Medina
|
1b3f678b50
|
feat: add Strategy 1 - 4-bit QLoRA with device_map=auto (distributed across GPUs)
|
2026-07-02 11:35:09 -04:00 |
|
Christian Medina
|
93ac7391dc
|
fix: clarify strategy naming + init distributed process group for FSDP
|
2026-07-02 11:30:34 -04:00 |
|
Christian Medina
|
8cabc0e986
|
fix: correct transformer_auto_wrap_policy API signature
|
2026-07-02 11:12:31 -04:00 |
|
Christian Medina
|
651965e844
|
feat: manually wrap model with FSDP on CPU before trainer
|
2026-07-02 11:03:41 -04:00 |
|
Christian Medina
|
b056ec0306
|
fix: force FSDP1 with string value to avoid FSDP2 gather spike
|
2026-07-02 10:54:17 -04:00 |
|
Christian Medina
|
f1e016c8af
|
fix: set gradient_checkpointing=False (FSDP handles it)
|
2026-07-02 09:38:51 -04:00 |
|
Christian Medina
|
361a58addc
|
fix: simplify FSDP config, comment out mixed_precision
|
2026-07-02 09:38:10 -04:00 |
|
Christian Medina
|
ab1988705b
|
fix: add use_orig_params=True for LoRA/PEFT compatibility
|
2026-07-02 09:37:37 -04:00 |
|
Christian Medina
|
b7d680966e
|
feat: restore FSDP with SHARD_GRAD_OP + sync_module_states
|
2026-07-02 09:36:08 -04:00 |
|
Christian Medina
|
7b39fc3a1b
|
fix: add --mixed_precision bf16 to accelerate launch
|
2026-07-02 09:26:37 -04:00 |
|
Christian Medina
|
fe7df7c92f
|
fix: disable FSDP, use standard accelerate data parallelism
|
2026-07-02 09:22:28 -04:00 |
|
Christian Medina
|
42d61c15b0
|
fix: disable sync_module_states + add mixed precision for FSDP
|
2026-07-02 09:14:18 -04:00 |
|
Christian Medina
|
4eb06a625e
|
fix: add bnb_4bit_quant_storage for FSDP sharding of 4-bit weights
|
2026-07-02 09:05:00 -04:00 |
|
Christian Medina
|
4c9522072e
|
fix: add BitsAndBytesConfig import + remove bf16 GPU strategy
|
2026-07-02 08:54:41 -04:00 |
|
Christian Medina
|
70a2bf0dd6
|
fix: use FSDP1 with sync_module_states to avoid GPU gather spike
|
2026-07-02 08:47:56 -04:00 |
|
Christian Medina
|
c607888d8f
|
fix: load QLoRA model to CPU first, FSDP shards later
|
2026-07-02 08:26:23 -04:00 |
|
Christian Medina
|
dbc97138e9
|
fix: use SHARD_GRAD_OP instead of FULL_SHARD for 2-GPU setup
|
2026-07-02 00:58:41 -04:00 |
|
Christian Medina
|
e3ea45fe96
|
fix: remove gradient_checkpointing (FSDP activation_checkpointing handles it)
|
2026-07-02 00:41:57 -04:00 |
|
Christian Medina
|
dfe85c3c82
|
fix: fsdp=True instead of list + add activation_checkpointing
|
2026-07-02 00:16:33 -04:00 |
|
Christian Medina
|
8747111671
|
fix: update config comment to reflect Ornith-35B
|
2026-07-02 00:11:10 -04:00 |
|
Christian Medina
|
d5ccd28e39
|
chore: remove temporary debugging code
|
2026-07-01 23:46:34 -04:00 |
|
Christian Medina
|
8a5edb25e0
|
fix: specify correct transformer layer class for FSDP auto-wrap
|
2026-07-01 23:45:05 -04:00 |
|
Christian Medina
|
05cbd7b6b2
|
show layer class names
|
2026-07-01 23:27:25 -04:00 |
|
Christian Medina
|
e64515e3ca
|
refactor: switch from DeepSpeed ZeRO-3 to FSDP for QLoRA compatibility
|
2026-07-01 22:33:06 -04:00 |
|
Christian Medina
|
a655040db2
|
feat: use DeepSpeed ZeRO-3 for proper model sharding across GPUs
|
2026-07-01 22:08:37 -04:00 |
|
Christian Medina
|
9d8c2b6cf0
|
refactor: cleaner loading strategies with error tracking
|
2026-07-01 21:38:38 -04:00 |
|
Christian Medina
|
815115e97a
|
feat: add BnB 4-bit quantization strategy + use bf16 model
|
2026-07-01 21:14:31 -04:00 |
|
Christian Medina
|
e56ae5e002
|
fix: restore all strategies with CPU-first as primary
|
2026-07-01 18:16:51 -04:00 |
|
Christian Medina
|
f18a9f0973
|
fix: use aggressive CPU loading for memory optimization
|
2026-07-01 18:00:55 -04:00 |
|
Christian Medina
|
83b45c2a5c
|
feat: add PYTORCH_CUDA_ALLOC_CONF for memory optimization
|
2026-07-01 17:01:28 -04:00 |
|
Christian Medina
|
19f405255c
|
remove auto device map in training strategies
|
2026-07-01 16:50:48 -04:00 |
|
Christian Medina
|
cc353b3dea
|
refactor: restructure project - scripts at root level
|
2026-07-01 16:39:54 -04:00 |
|
Christian Medina
|
293f9caf65
|
fix: use correct repo_root path for dataset files
|
2026-07-01 16:35:47 -04:00 |
|
Christian Medina
|
1095032801
|
feat: add ai none to unload AI model before training
|
2026-07-01 16:24:30 -04:00 |
|
Christian Medina
|
79aa1e4061
|
fix: don't reinstall PyTorch (already installed)
|
2026-07-01 16:22:55 -04:00 |
|
Christian Medina
|
3b771c0db5
|
fix: add venv setup and proper pip install
|
2026-07-01 16:22:09 -04:00 |
|
Christian Medina
|
ff0d6e8736
|
fix: use correct path for training script
|
2026-07-01 16:21:47 -04:00 |
|
Christian Medina
|
5eae235180
|
executable permissions
|
2026-07-01 16:06:52 -04:00 |
|
Christian Medina
|
7e13680e27
|
additional fixes
|
2026-07-01 16:05:52 -04:00 |
|
Christian Medina
|
088521094e
|
feat: add 4 loading strategies (no BnB for already-quantized model)
|
2026-07-01 15:57:42 -04:00 |
|
Christian Medina
|
3089f27901
|
fix: load 4-bit model AS-IS without BnB (already quantized)
|
2026-07-01 15:56:00 -04:00 |
|
Christian Medina
|
1197235b4d
|
fix: skip prepare_model_for_kbit_training to avoid OOM
|
2026-07-01 15:18:48 -04:00 |
|
Christian Medina
|
d9fd079686
|
fix: add accelerate CPU offload to bf16 strategies
|
2026-07-01 15:14:00 -04:00 |
|
Christian Medina
|
4ed80ad3a3
|
feat: restore all 6 loading strategies
|
2026-07-01 15:11:45 -04:00 |
|
Christian Medina
|
313b44381f
|
fix: remove DeepSpeed, use torchrun without distributed training
|
2026-07-01 15:10:54 -04:00 |
|
Christian Medina
|
7de54746a2
|
feat: add 6 loading variants with different offload strategies
|
2026-07-01 14:45:29 -04:00 |
|
Christian Medina
|
22d1bfd573
|
fix: use 4-bit model first, then bf16 with CPU offload
|
2026-07-01 14:43:30 -04:00 |
|
Christian Medina
|
3fa2cfe4cc
|
fix: use DeepSpeed ZeRO-3 directly (skip CPU loading)
|
2026-07-01 14:41:48 -04:00 |
|
Christian Medina
|
3760aaf316
|
fix: rename output dir to ornith-35b-lora
|
2026-07-01 14:19:43 -04:00 |
|
Christian Medina
|
4287e42148
|
fix: remove screen, run training directly
|
2026-07-01 14:16:05 -04:00 |
|
Christian Medina
|
3bd51ec7db
|
feat: add screen support to bypass cgroup limits
|
2026-07-01 14:09:34 -04:00 |
|
Christian Medina
|
ce758f7eb5
|
feat: improve loading strategies with low_cpu_mem_usage and proper DeepSpeed config
|
2026-07-01 14:00:41 -04:00 |
|
Christian Medina
|
2d13e63c25
|
feat: add 4 loading strategies with automatic fallback
|
2026-07-01 13:55:23 -04:00 |
|
Christian Medina
|
b300235be0
|
fix: load 4-bit model AS-IS without additional quantization
|
2026-07-01 13:53:43 -04:00 |
|
Christian Medina
|
e967bd74f1
|
fix: skip quantization, use DeepSpeed ZeRO-3 CPU offload directly
|
2026-07-01 13:53:06 -04:00 |
|
Christian Medina
|
7539e39a93
|
fix: use accelerate load_checkpoint_and_dispatch for proper CPU offload
|
2026-07-01 13:39:51 -04:00 |
|
Christian Medina
|
528e7002ef
|
feat: add 8-bit CPU offload fallback between 4-bit and bf16
|
2026-07-01 13:32:04 -04:00 |
|
Christian Medina
|
59ddfb45ed
|
feat: add fallback to bf16 with CPU offload if 4-bit fails
|
2026-07-01 13:29:23 -04:00 |
|
Christian Medina
|
b41d99f956
|
fix: remove broken quantization_config before loading
|
2026-07-01 13:12:57 -04:00 |
|
Christian Medina
|
1228a1f38b
|
fix: load model first then quantize with BnB
|
2026-07-01 13:11:14 -04:00 |
|
Christian Medina
|
14042dbb44
|
fix: use pre-quantized 4bit model
|
2026-07-01 12:35:19 -04:00 |
|
Christian Medina
|
f784d06f0a
|
fix: load QLoRA directly to GPU (skip CPU)
|
2026-07-01 12:33:21 -04:00 |
|
Christian Medina
|
9471d5cc77
|
fix: cast learning_rate to float (YAML parses 2e-4 as string)
|
2026-07-01 11:19:00 -04:00 |
|
Christian Medina
|
2fa52cd9d4
|
fix: load QLoRA model to CPU first, DeepSpeed distributes
|
2026-07-01 09:37:20 -04:00 |
|
Christian Medina
|
2fa5173c56
|
feat: add prepare_model_for_kbit_training and update LoRA r=64 alpha=128
|
2026-07-01 07:48:44 -04:00 |
|
Christian Medina
|
357c9fa781
|
fix: use standard QLoRA pattern with BitsAndBytesConfig
|
2026-07-01 07:47:42 -04:00 |
|
Christian Medina
|
da5eb3abed
|
feat: implement QLoRA with 4-bit BitsAndBytes quantization
|
2026-07-01 07:45:33 -04:00 |
|
Christian Medina
|
f0ee6bc9a2
|
fix: torch_dtype -> dtype (fix deprecation warning)
|
2026-07-01 07:38:36 -04:00 |
|
Christian Medina
|
f651ca030c
|
fix: convert warmup_ratio to warmup_steps (fix deprecation warning)
|
2026-07-01 07:37:47 -04:00 |
|
Christian Medina
|
f47944ee75
|
fix: remove quantization_config after loading to bypass SFTTrainer check
|
2026-07-01 07:36:38 -04:00 |
|
Christian Medina
|
86b401c12f
|
fix: convert FP8 to bf16 for PEFT compatibility
|
2026-07-01 07:30:10 -04:00 |
|
Christian Medina
|
6f631f00ae
|
fix: use FP8 model with 32GB CPU offload
|
2026-07-01 07:29:32 -04:00 |
|
Christian Medina
|
9d92531c55
|
fix: use Ornith-1.0-35B bf16 base model (PEFT compatible)
|
2026-07-01 07:28:21 -04:00 |
|
Christian Medina
|
05ace2d38f
|
fix: load NVFP4 as bf16 by overriding quantization config
|
2026-07-01 07:27:13 -04:00 |
|
Christian Medina
|
be7b589c6b
|
fix: use BitsAndBytes INT4 for PEFT compatibility
|
2026-07-01 00:34:34 -04:00 |
|
Christian Medina
|
d7301d37dd
|
fix: use float16 dtype to preserve NVFP4 quantization
|
2026-07-01 00:08:42 -04:00 |
|
Christian Medina
|
af09d47c6f
|
fix: load model to CPU first, let DeepSpeed distribute
|
2026-07-01 00:03:24 -04:00 |
|
Christian Medina
|
2c1251788d
|
fix: add compressed-tensors to training dependencies
|
2026-06-30 23:56:59 -04:00 |
|
Christian Medina
|
905c2ba30d
|
fix: preserve NVFP4 quantization, don't force bf16
|
2026-06-30 23:50:27 -04:00 |
|
Christian Medina
|
114e25fcfb
|
fix: use device_map=auto to distribute model across GPUs
|
2026-06-30 23:45:57 -04:00 |
|
Christian Medina
|
8d4cdca3bb
|
feat: use Ornith-1.0-35B-NVFP4 for training
|
2026-06-30 23:34:33 -04:00 |
|
Christian Medina
|
5713bf0285
|
fix: add CPU offload back to DeepSpeed ZeRO-3
|
2026-06-30 22:47:19 -04:00 |
|
Christian Medina
|
3bb7328e2a
|
fix: disable quantization config for fp8 model training
|
2026-06-30 22:41:28 -04:00 |
|
Christian Medina
|
a8c9cd682e
|
fix: simplify DeepSpeed config, remove offload settings
|
2026-06-30 22:36:49 -04:00 |
|
Christian Medina
|
b760ba09e7
|
fix: use CUDA 13.0 nightly for RTX 5090 support
|
2026-06-30 22:32:14 -04:00 |
|
Christian Medina
|
9cbf6549e5
|
fix: use output_file parameter in prepare_dataset
|
2026-06-30 21:54:05 -04:00 |
|
Christian Medina
|
ed0b1eddf8
|
fix: use messages format for SFTTrainer (system/user/assistant)
|
2026-06-30 21:05:48 -04:00 |
|
Christian Medina
|
d8ad31fd03
|
fix: remove max_seq_length from SFTTrainer (newer TRL API)
|
2026-06-30 20:59:24 -04:00 |
|
Christian Medina
|
8ad87f0e7e
|
fix: remove dataset_text_field (dataset pre-formatted)
|
2026-06-30 20:48:07 -04:00 |
|