model: support youtu-vl model (#18479)
* Support Youtu-VL Model * merge code * fix bug * revert qwen2 code & support rsplit in minja.hpp * update warm info * fix annotation * u * revert minja.hpp * fix * Do not write routed_scaling_factor to gguf when routed_scaling_factor is None * fix expert_weights_scale * LGTM after whitespace fixes * fix * fix * fix * layers to layer_index * enum fix --------- Co-authored-by: Xuan-Son Nguyen <son@huggingface.co> Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@scala.com>
This commit is contained in:
co-authored by
Xuan-Son Nguyen
Sigbjørn Skjæret
parent
3ccccc83f7
commit
ced765be44
@@ -1129,11 +1129,40 @@ class GGUFWriter:
|
||||
self.add_uint32(Keys.ClipVision.Projector.SCALE_FACTOR, value)
|
||||
|
||||
def add_vision_n_wa_pattern(self, value: int) -> None:
|
||||
"""Add window attention pattern interval for vision models.
|
||||
|
||||
This defines the pattern interval for window attention vs full attention layers.
|
||||
For example, if n_wa_pattern=4, then layers 3, 7, 11, ... use full attention,
|
||||
while other layers use window attention.
|
||||
|
||||
Used by models like Qwen2.5-VL where full attention layers follow a regular pattern.
|
||||
"""
|
||||
self.add_uint32(Keys.ClipVision.N_WA_PATTERN, value)
|
||||
|
||||
def add_vision_wa_layer_indexes(self, layers: Sequence[int]) -> None:
|
||||
"""Add explicit layer indexes that use full attention in vision models.
|
||||
|
||||
This specifies the exact layer indices (0-based) that should use full attention
|
||||
instead of window attention. All other layers will use window attention.
|
||||
|
||||
Args:
|
||||
layers: List of layer indices that use full attention (e.g., [3, 7, 11, 15])
|
||||
|
||||
Used by models like YoutuVL where full attention layers are explicitly specified
|
||||
rather than following a regular pattern.
|
||||
|
||||
Difference from add_vision_n_wa_pattern:
|
||||
- n_wa_pattern: Defines a regular interval pattern (every Nth layer uses full attention)
|
||||
- wa_layer_indexes: Explicitly lists which layers use full attention (irregular pattern)
|
||||
"""
|
||||
self.add_array(Keys.ClipVision.WA_LAYER_INDEXES, layers)
|
||||
|
||||
def add_vision_is_deepstack_layers(self, layers: Sequence[bool]) -> None:
|
||||
self.add_array(Keys.ClipVision.IS_DEEPSTACK_LAYERS, layers)
|
||||
|
||||
def add_vision_window_size(self, value: int) -> None:
|
||||
self.add_uint32(Keys.ClipVision.WINDOW_SIZE, value)
|
||||
|
||||
# audio models
|
||||
|
||||
def add_audio_projection_dim(self, value: int) -> None:
|
||||
|
||||
Reference in New Issue
Block a user