Longformer: Process 4096-Token Documents With Sliding Window Attention
Longformer extends BERT to 4096 tokens using a combination of local sliding window attention and global attention, making it practical for document classification, Q&A, and NER on long-form text.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
Standard transformer attention is quadratic in sequence length - doubling the sequence length quadruples computation. BERT's 512-token limit means long documents must be chunked, losing inter-chunk context. A question that spans two chunks, or a document summary that requires reading the conclusion before understanding the introduction, breaks with naive chunking.
Longformer from AllenAI solves this with two attention mechanisms applied together:
- Sliding window attention: Each token attends to
window_sizeneighbors on each side. Local context is preserved efficiently at O(n × w) cost. - Global attention: Selected tokens (like
[CLS]or question tokens) attend to all other tokens and are attended to by all other tokens. Global tokens gather document-level context.
This combination maintains O(n) complexity while preserving the ability to reason across the full document.
When to Use Longformer vs Chunking Strategies
Use Longformer when:
- Questions or labels depend on evidence scattered across the document
- You need document-level representations (not sentence-level)
- Documents are 600-4000 tokens consistently
Use chunking when:
- Documents are 4000+ tokens (Longformer's 4096 limit still applies)
- Questions can always be answered from a local passage
- Throughput matters more than accuracy (chunking + retrieval is faster)
The sweet spot is 800-2500 token documents - long enough that BERT fails, short enough that Longformer handles cleanly.
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Document Classification
from transformers import LongformerForSequenceClassification, LongformerTokenizerFast
import torch
tokenizer = LongformerTokenizerFast.from_pretrained("allenai/longformer-base-4096")
model = LongformerForSequenceClassification.from_pretrained(
"allenai/longformer-base-4096",
num_labels=4
)
long_document = "..." * 1000 # 3000+ word document
inputs = tokenizer(
long_document,
return_tensors="pt",
max_length=4096,
truncation=True,
padding="max_length"
)
# Global attention on [CLS] token
global_attention_mask = torch.zeros_like(inputs["input_ids"])
global_attention_mask[:, 0] = 1
outputs = model(**inputs, global_attention_mask=global_attention_mask)
logits = outputs.logits
predicted_class = logits.argmax(dim=-1).item()
The HuggingFace Longformer-large-4096 is the highest-quality variant; base is 3x faster.
Document Q&A With LongformerForQuestionAnswering
from transformers import LongformerForQuestionAnswering, LongformerTokenizerFast
model = LongformerForQuestionAnswering.from_pretrained("allenai/longformer-large-4096-finetuned-triviaqa")
tokenizer = LongformerTokenizerFast.from_pretrained("allenai/longformer-large-4096-finetuned-triviaqa")
question = "What is the main argument in section 3?"
document = "...full document text..."
encoding = tokenizer(question, document, return_tensors="pt", max_length=4096, truncation=True)
# Global attention on question tokens
sequence_ids = encoding.sequence_ids(0)
global_attention = [1 if t == 0 else 0 for t in sequence_ids]
encoding["global_attention_mask"] = torch.tensor([global_attention])
outputs = model(**encoding)
start = outputs.start_logits.argmax()
end = outputs.end_logits.argmax()
answer = tokenizer.decode(encoding["input_ids"][0][start:end+1])
print(answer)
Comparison to Big Bird
Big Bird (Google) solves the same long-sequence problem with a different attention pattern: random attention + sliding window + global. Both achieve similar accuracy on long-document benchmarks. Longformer is more widely deployed and has better HuggingFace ecosystem support; Big Bird's random attention may provide marginal advantages on extremely long documents (>4096 tokens with extended variants).

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
Machine Learning: Complete Guide for Software Developers
Learn machine learning as a software developer with this complete guide covering Python, algorithms, mathematics, projects, deep learning, and a practical roadmap.
ONNX: Export Any ML Model and Run It Anywhere
ONNX (Open Neural Network Exchange) is the universal model format - export from PyTorch, scikit-learn, or HuggingFace and run 3x faster inference with ONNX Runtime on CPU or GPU.
Gradient Descent Explained: How Machine Learning Models Actually Learn
Gradient descent is the engine behind every modern ML model. Here is how it works, why learning rate matters, and when to use Adam over SGD.
// discussion
Comments