Dark Hacker News
new
|
best
|
ask
|
show
|
jobs
1 points
by
g023
26 days ago
undefined | Dark Hacker News
g023
26 days ago
|
next
[−]
A self-contained CUDA inference engine for LiquidAI/LFM2.5-8B-A1B (hybrid conv + GQA-attention MoE, 8.5B params, 1B active) targeting a single RTX 3060 (12 GB) using flash-decoding. MIT license.