-
Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Paper • 2402.19427 • Published • 52 -
Simple linear attention language models balance the recall-throughput tradeoff
Paper • 2402.18668 • Published • 18 -
ChunkAttention: Efficient Self-Attention with Prefix-Aware KV Cache and Two-Phase Partition
Paper • 2402.15220 • Published • 19 -
Linear Transformers are Versatile In-Context Learners
Paper • 2402.14180 • Published • 6
Kiran Kamble
kiranr
AI & ML interests
nlp,llm
Recent Activity
liked
a dataset
9 days ago
Gryphe/Sonnet3.5-SlimOrcaDedupCleaned
liked
a dataset
9 days ago
Magpie-Align/Magpie-Llama-3.1-Pro-MT-300K-Filtered
liked
a dataset
9 days ago
virattt/financial-qa-10K
Organizations
Collections
1
models
1
datasets
None public yet