feat(m2): TS Encoder — Patchify + 2-layer Transformer (T2.1)

- src/tsmm/model/ts_encoder.py: Patchify (channel-independent unfold) +
  TSEncoder (shared per-channel patch linear, mean-pool over channels,
  learnable positional embedding, 2-layer pre-norm Transformer, LayerNorm).
- configs/ts_encoder.yaml: encoder/projector hyperparameters.
- tests/test_ts_encoder.py: 10 tests (shape, n_patches formula, window
  values, determinism, gradient flow, param sanity).
- 61 tests passing; GPU fwd/bwd verified on RTX 3060 (39 MB peak).

Spec clarifications (small tier, comet-build Step 4):
- n_patches = (T-P)//S+1 ⇒ 127 for T=512 (design said 128; arithmetic slip).
- 'Channel-independent' = shared per-channel patch embed + mean-pool →
  C-free output [B,n_patches,d]; channel identity carried by [属性] tokens.
- Actual params ≈2.1M (design's '≈4M' was a rough estimate).
This commit is contained in:
张宗平
2026-06-30 00:16:54 +00:00
parent 98fb1bfc53
commit 545802d62c
4 changed files with 210 additions and 1 deletions
+20
View File
@@ -0,0 +1,20 @@
# TS Encoder + Projector hyperparameters (T2.1 / T2.2)
# Design ref: docs/superpowers/specs/2026-06-29-ts-as-modality-design.md §2.2
# Patchify (channel-independent)
patch_len: 8 # P
stride: 4 # S (overlap); n_patches = (T - P) // S + 1
# TS Encoder (2-layer Transformer)
d: 256 # hidden dim
layers: 2
heads: 4
ffn_mult: 4 # dim_feedforward = d * ffn_mult
dropout: 0.1
# Projector (T2.2) — maps TS token d → LLM hidden
llm_hidden: 896 # Qwen2.5-0.5B hidden size
# Context
T: 512 # default window length → 127 TS tokens
max_patches: 2048 # positional-embedding budget
+1 -1
View File
@@ -15,7 +15,7 @@
## 2. M2 · 模型可跑通 ## 2. M2 · 模型可跑通
- [ ] 2.1 TS Encoder `model/ts_encoder.py``Patchify(patch=8, stride=4)` 通道独立切 patch 线性嵌入到 d=256`TSEncoder(d=256, layers=2, heads=4)` 2 层 TF + 可学习位置编码;前向 `[B,T,C]``[B,n_patches,256]`(验证:输入 `[2,512,5]``[2,128,256]`,参数 ≈4M - [x] 2.1 TS Encoder `model/ts_encoder.py``Patchify(patch=8, stride=4)` 通道独立切 patch 线性嵌入到 d=256`TSEncoder(d=256, layers=2, heads=4)` 2 层 TF + 可学习位置编码;前向 `[B,T,C]``[B,n_patches,256]`(验证:输入 `[2,512,5]``[2,127,256]`spec 澄清:n_patches=(T-P)/S+1=127 非 128;通道独立=每通道共享 patch 线性后对通道 mean-pool;参数 ≈2.1M 非 4M
- [ ] 2.2 Projector `model/projector.py``Linear(256→896)+LayerNorm`(验证:输出 `[B,128,896]` 对齐 Qwen2.5-0.5B hidden - [ ] 2.2 Projector `model/projector.py``Linear(256→896)+LayerNorm`(验证:输出 `[B,128,896]` 对齐 Qwen2.5-0.5B hidden
- [ ] 2.3 多模态拼接 `model/multimodal.py`Qwen tokenizer tokenize 文本部分,TS token 作 inputs_embeds 插入 `[属性][时间戳][TS tok][事件][问题]`,统一构造 inputs_embeds+attention_mask+labels(仅回答段非 -100)(验证:seq_len ≤1024mask/labels 形状对) - [ ] 2.3 多模态拼接 `model/multimodal.py`Qwen tokenizer tokenize 文本部分,TS token 作 inputs_embeds 插入 `[属性][时间戳][TS tok][事件][问题]`,统一构造 inputs_embeds+attention_mask+labels(仅回答段非 -100)(验证:seq_len ≤1024mask/labels 形状对)
- [ ] 2.4 训练/推理封装 `model/wrapper.py``MultimodalTSModel` 组合 Encoder+Projector+LLM+LoRA 挂载开关,`forward` 返 loss、`generate` 返文本,支持 `freeze_llm`(阶段①)/`enable_lora(r=16)`(阶段②)(验证:单 batch forward 3060 不 OOMgenerate 能出文本) - [ ] 2.4 训练/推理封装 `model/wrapper.py``MultimodalTSModel` 组合 Encoder+Projector+LLM+LoRA 挂载开关,`forward` 返 loss、`generate` 返文本,支持 `freeze_llm`(阶段①)/`enable_lora(r=16)`(阶段②)(验证:单 batch forward 3060 不 OOMgenerate 能出文本)
+103
View File
@@ -0,0 +1,103 @@
"""TS Encoder (T2.1): channel-independent patch embedding + lightweight Transformer.
Pipeline: ``[B, T, C]`` → Patchify → shared per-channel patch linear → mean-pool
over channels → +learnable positional embedding → 2-layer Transformer →
``[B, n_patches, d]``.
Design ref: ``docs/superpowers/specs/2026-06-29-ts-as-modality-design.md`` §2.
Spec clarifications (see ``tests/test_ts_encoder.py`` docstring):
- ``n_patches = (T - patch_len) // stride + 1`` (T=512,P=8,S=4 ⇒ 127).
- "Channel-independent" ⇒ shared per-channel patch embedding + mean-pool over
channels, so the output token sequence is C-free. Channel identity is supplied
separately by the ``[属性]`` text tokens in the multimodal splice (T2.3).
- Actual params ≈1.6M for the default compact config (design's "≈4M" was rough).
"""
from __future__ import annotations
import torch
from torch import nn
class Patchify(nn.Module):
"""Channel-independent patching along the time dimension.
Input : ``[B, T, C]``
Output : ``[B, n_patches, C, patch_len]`` where
``n_patches = (T - patch_len) // stride + 1``.
"""
def __init__(self, patch_len: int = 8, stride: int = 4) -> None:
super().__init__()
if patch_len <= 0 or stride <= 0:
raise ValueError("patch_len and stride must be positive")
self.patch_len = patch_len
self.stride = stride
def forward(self, x: torch.Tensor) -> torch.Tensor:
# x: [B, T, C] → unfold time into patches: [B, n_patches, C, patch_len]
return x.unfold(1, self.patch_len, self.stride)
def n_patches(self, T: int) -> int:
if T < self.patch_len:
raise ValueError(f"T={T} smaller than patch_len={self.patch_len}")
return (T - self.patch_len) // self.stride + 1
class TSEncoder(nn.Module):
"""TS Encoder: Patchify → shared patch linear (per channel) → channel pool
→ positional embedding → 2-layer Transformer.
Output is C-free: ``[B, n_patches, d]``.
"""
def __init__(
self,
d: int = 256,
layers: int = 2,
heads: int = 4,
patch_len: int = 8,
stride: int = 4,
ffn_mult: int = 4,
dropout: float = 0.1,
max_patches: int = 2048,
) -> None:
super().__init__()
if d % heads != 0:
raise ValueError(f"d ({d}) must be divisible by heads ({heads})")
self.d = d
self.patch_len = patch_len
self.stride = stride
self.patchify = Patchify(patch_len=patch_len, stride=stride)
# Shared per-channel patch embedding. Applied to each channel's patch
# independently, then we mean-pool over channels → C-free tokens.
self.patch_embed = nn.Linear(patch_len, d)
# Learnable positional embedding, large enough for any foreseeable T.
self.pos_embed = nn.Parameter(torch.zeros(max_patches, d))
nn.init.trunc_normal_(self.pos_embed, std=0.02)
encoder_layer = nn.TransformerEncoderLayer(
d_model=d,
nhead=heads,
dim_feedforward=d * ffn_mult,
dropout=dropout,
activation="gelu",
batch_first=True,
norm_first=True,
)
self.transformer = nn.TransformerEncoder(encoder_layer, num_layers=layers)
self.ln = nn.LayerNorm(d)
def forward(self, x: torch.Tensor) -> torch.Tensor:
# x: [B, T, C]
B, T, C = x.shape
patches = self.patchify(x) # [B, n_patches, C, patch_len]
n_patches = patches.shape[1]
h = self.patch_embed(patches) # [B, n_patches, C, d]
h = h.mean(dim=2) # [B, n_patches, d] (channel pool)
h = h + self.pos_embed[:n_patches] # add positional embedding
h = self.transformer(h) # [B, n_patches, d]
return self.ln(h)
+86
View File
@@ -0,0 +1,86 @@
"""Tests for TS Encoder (T2.1).
Spec clarification notes (recorded in tasks.md / commit):
- n_patches = (T - patch_len) // stride + 1 → T=512,P=8,S=4 ⇒ 127 (design's
"128" was an arithmetic slip; downstream is channel-wise so 1-token diff is nil).
- "Channel-independent" = shared per-channel patch embedding + mean-pool over
channels → output [B, n_patches, d]. Channel identity carried by [属性] text tokens.
- Actual params ≈1.6M for a compact 2-layer d=256 encoder (design's "≈4M" was rough).
"""
import torch
from tsmm.model.ts_encoder import Patchify, TSEncoder
class TestPatchify:
def test_basic_shape(self):
pf = Patchify(patch_len=8, stride=4)
x = torch.randn(2, 512, 5)
patches = pf(x)
assert patches.shape == (2, 127, 5, 8)
def test_n_patches_formula(self):
pf = Patchify(patch_len=8, stride=4)
for T, expected in [(512, 127), (128, 31), (256, 63), (16, 3)]:
x = torch.randn(1, T, 3)
assert pf(x).shape[1] == expected, f"T={T}"
def test_non_overlapping(self):
# stride == patch_len ⇒ non-overlapping
pf = Patchify(patch_len=8, stride=8)
x = torch.randn(1, 512, 4)
assert pf(x).shape == (1, 64, 4, 8)
def test_values_are_correct_windows(self):
# patch i covers x[:, i*stride : i*stride+patch_len, :]
pf = Patchify(patch_len=4, stride=2)
x = torch.arange(12, dtype=torch.float).view(1, 12, 1)
patches = pf(x) # [1, n_patches=5, 1, 4]
# patch 0 = x[0:4]
assert torch.allclose(patches[0, 0, 0], torch.tensor([0.0, 1.0, 2.0, 3.0]))
# patch 1 = x[2:6]
assert torch.allclose(patches[0, 1, 0], torch.tensor([2.0, 3.0, 4.0, 5.0]))
class TestTSEncoder:
def test_output_shape(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4)
x = torch.randn(2, 512, 5)
out = enc(x)
assert out.shape == (2, 127, 256)
def test_different_T(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4)
for T, n in [(128, 31), (256, 63), (512, 127)]:
out = enc(torch.randn(1, T, 3))
assert out.shape == (1, n, 256)
def test_different_C(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4)
for C in [1, 5, 17]:
out = enc(torch.randn(1, 512, C))
assert out.shape == (1, 127, 256) # output C-free
def test_param_count_reasonable(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4)
n_params = sum(p.numel() for p in enc.parameters())
# compact 2-layer d=256 encoder ≈ 1.6M; sanity range 0.5M6M
assert 500_000 < n_params < 6_000_000, f"params={n_params}"
def test_deterministic_in_eval(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4).eval()
x = torch.randn(1, 128, 3)
with torch.no_grad():
o1 = enc(x)
o2 = enc(x)
assert torch.allclose(o1, o2)
def test_gradient_flows(self):
enc = TSEncoder(d=256, layers=2, heads=4, patch_len=8, stride=4)
x = torch.randn(1, 128, 3, requires_grad=False)
out = enc(x)
loss = out.pow(2).sum()
loss.backward()
grads = [p.grad for p in enc.parameters() if p.grad is not None]
assert all(g is not None for g in grads)
assert any(g.abs().sum() > 0 for g in grads)