Free tools Windows power users keep installed
One-click scans. No signup required.
This guide builds a small Transformer neural network from scratch using PyTorch. You will implement attention, multi-head self-attention, positional embeddings, decoder blocks, training, testing, and text generation. It does not cover winding or wiring a mains-voltage electrical transformer. That is a separate electrical-engineering project with serious shock, fire, and insulation hazards.
The finished model is an educational, inspectable decoder-only language model. It can learn patterns from a small text corpus and generate text, but it will not reproduce ChatGPT, Llama, or another production large language model.
What you are building
A Transformer is a neural-network architecture introduced in the 2017 paper “Attention Is All You Need”. Instead of relying primarily on recurrence or convolution to process a sequence, it uses learned, content-dependent interactions between positions.
There are several kinds of Transformer project:
| Project | Typical use |
|---|---|
| Attention from scratch | Understand queries, keys, values, and attention weights |
| Encoder-only Transformer | Classification, embeddings, and BERT-style tasks |
| Decoder-only Transformer | Autoregressive text generation and GPT-style models |
| Encoder–decoder Transformer | Translation and other sequence-to-sequence tasks |
| Vision Transformer | Image classification using image patches as tokens |
This implementation uses the decoder-only design because it provides a clear result: given previous tokens, the model predicts the next token.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
- POWERED BY RTX 5070 12GB + RYZEN 7 9700X - The GeForce RTX 5070 12GB GDDR7 graphics card pairs with an 8-core AMD Ryzen 7 9700X processor to drive smooth 1440p and 4K gameplay, giving this gaming PC the headroom for modern titles, streaming, and creative work.
- 32GB DDR5 6000MHz MEMORY & 1TB NVMe SSD - 32GB of high-speed DDR5 memory and a 1TB PCIe 4.0 NVMe solid state drive deliver quick load times, smooth multitasking, and generous storage, keeping this prebuilt gaming desktop responsive under heavy workloads.
- BUILT-IN 11.3-INCH Smart DISPLAY - An integrated smart screen shows real-time CPU and GPU temperatures, usage, and weather while you play, adding a distinctive and functional touch to your battlestation.
- 850W 80+ GOLD POWER SUPPLY, 360MM LIQUID COOLING & WiFi 7 - An 850W 80 Plus Gold certified power supply provides stable, efficient power with headroom for future upgrades, while a 360mm AIO liquid cooler, WiFi 7, and an ARGB mid-tower case keep the Ryzen 7 CPU cool and connected in a clean build.
- READY TO PLAY OUT OF THE BOX - Arrives fully assembled and tested with Windows 11 Home pre-installed, so your prebuilt gaming computer is ready to set up in minutes. Assembled in the USA, and backed by a one-year limited warranty and lifetime free technical support.
How a Transformer represents text
Text must first become numbers. A tokenizer divides text into tokens, then maps each token to an integer ID. Tokens may be characters, words, subwords, or special markers such as beginning-of-sequence, end-of-sequence, and padding tokens.
For the first implementation, character-level tokenization keeps the data pipeline transparent:
text → characters → integer IDs → embeddings → Transformer blocks → logits
Character tokenization has a small vocabulary and requires few dependencies, but it produces longer sequences. Production systems commonly use subword tokenization because word-level vocabularies become very large while character-level sequences can be inefficient.
Install the environment
Use a virtual environment and install the basic dependencies:
python -m venv .venv
source .venv/bin/activate # macOS/Linux
.venvScriptsactivate # Windows PowerShell
python -m pip install --upgrade pip
pip install torch numpy matplotlib tqdm
Check the local installation rather than assuming a particular version:
python --version
python -c "import torch; print(torch.__version__)"
python -c "import torch; print(torch.cuda.is_available())"
python -c "import torch; print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'CPU')"
CPU execution is sufficient for shape tests and tiny experiments. Google Colab is another practical option for small educational models. Hardware availability, session duration, CUDA support, and runtime vary, so do not treat one reported GPU workload as a general training estimate.
The architecture
token IDs
↓
token embeddings + positional embeddings
↓
pre-normalized decoder block
├── causal multi-head self-attention
├── residual connection
├── feed-forward network
└── residual connection
↓
repeat N times
↓
final layer normalization
↓
linear vocabulary projection
↓
next-token logits
For a batch of sequences, use these symbols:
B: batch sizeT: sequence lengthC: model dimensionH: number of attention headsD: head dimension, whereD = C / HV: vocabulary size
Prepare a small corpus
Start with a small, legally usable text file. Build a character vocabulary and conversion functions:
Rank #2
- Powerful Processor: AMD Ryzen 5 5600GT 3.6GHz (4.6GHz Turbo) 6-Core 12-Thread processor brings faster response time to easily handle multi-threaded tasks
- Motherboard Specification: MSI A520M-A PRO motherboard provides reliable performance and expandability for your computing needs
- Integrated Graphics: AMD Radeon Vega Graphics (CPU Integration) enables you to play 1080P mainstream games at quality frame rates
- Memory and Storage: 16GB DDR4 3200MHz RAM paired with 1TB M.2 NVMe PCIe SSD for fast multitasking and quick data access
- Power Supply: 550W 80PLUS Bronze certified power supply ensures stable and energy-efficient operation
from pathlib import Path
text = Path("corpus.txt").read_text(encoding="utf-8")
chars = sorted(set(text))
vocab_size = len(chars)
stoi = {ch: i for i, ch in enumerate(chars)}
itos = {i: ch for ch, i in stoi.items()}
encode = lambda s: [stoi[ch] for ch in s]
decode = lambda ids: "".join(itos[i] for i in ids)
ids = encode(text)
Split the underlying text stream before creating overlapping windows. If you create all windows first, nearly identical windows can appear in both training and validation data.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Create shifted training examples
A decoder-only language model learns next-token prediction. For:
Input: The cat sat
Target: cat sat on
the target is shifted one position to the left. The dataset returns a sequence of inputs and the same sequence shifted by one token:
import torch
from torch.utils.data import Dataset
class TextDataset(Dataset):
def __init__(self, ids, context_length):
self.ids = ids
self.context_length = context_length
def __len__(self):
return len(self.ids) - self.context_length
def __getitem__(self, index):
chunk = self.ids[index:index + self.context_length + 1]
x = torch.tensor(chunk[:-1], dtype=torch.long)
y = torch.tensor(chunk[1:], dtype=torch.long)
return x, y
Every token ID must be an integer tensor with dtype torch.long. The sequence length must not exceed the model's configured context length.
Embeddings and positional information
An embedding table turns each token ID into a learned vector:
token_embedding = nn.Embedding(vocab_size, d_model)
Self-attention alone does not know order. Without positional information, the same token vectors would represent the same set regardless of sequence order.
The original Transformer used fixed sinusoidal positional encodings. This implementation uses learned positional embeddings because they are simple for a small GPT-style model:
Rank #3
- Intel Core i7 14700F, NVIDIA GeForce RTX 5070 12GB, 32GB DDR5 RGB 4800MHz 16x2 1TB NVMe SSD, WIFI Ready, Windows 11 Home
- Connectivity: 6 x USB 3.1 | 1x RJ-45 Network Ethernet 10/100/1000 | Audio: On board audio
- Special Add-Ons: Tempered Glass RGB Gaming Case | 802.11AC Wi-Fi Included | 16 Color RGB Lighting Case | Free iBUYPOWER Gaming Keyboard & RGB Gaming Mouse | No Bloatware | AI Workstation PC ready
self.position_embedding = nn.Embedding(context_length, d_model)
positions = torch.arange(seq_len, device=x.device)
x = token_embedding(token_ids) + self.position_embedding(positions)
Learned positions are tied to the configured maximum context length. Sinusoidal and other positional methods are valid alternatives; neither should be presented as universally best.
Implement scaled dot-product attention
Attention uses queries, keys, and values. A query is compared with keys to produce weights, and those weights form a mixture of value vectors:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Attention(Q,K,V) = softmax((QKᵀ / √dₖ) + M)V
The division by √dₖ matters. As the key dimension grows, unscaled dot products tend to grow in magnitude. Very large softmax inputs produce excessively sharp distributions and weaker useful gradients.
import math
import torch
import torch.nn.functional as F
def scaled_dot_product_attention(q, k, v, mask=None, dropout=None):
d_k = q.size(-1)
scores = q @ k.transpose(-2, -1)
scores = scores / math.sqrt(d_k)
if mask is not None:
scores = scores.masked_fill(mask == 0, float("-inf"))
weights = torch.softmax(scores, dim=-1)
if dropout is not None:
weights = dropout(weights)
return weights @ v, weights
With multiple heads, the expected shapes are:
q: (B, H, T, D)
k: (B, H, T, D)
v: (B, H, T, D)
scores: (B, H, T, T)
output: (B, H, T, D)
Add causal multi-head self-attention
Each head attends within a smaller representation subspace. If d_model = 128 and num_heads = 4, each head has dimension 32. The model dimension must divide evenly by the number of heads.
import torch.nn as nn
class MultiHeadSelfAttention(nn.Module):
def __init__(self, d_model, num_heads, dropout=0.0):
super().__init__()
assert d_model % num_heads == 0
self.num_heads = num_heads
self.head_dim = d_model // num_heads
self.q_proj = nn.Linear(d_model, d_model)
self.k_proj = nn.Linear(d_model, d_model)
self.v_proj = nn.Linear(d_model, d_model)
self.out_proj = nn.Linear(d_model, d_model)
self.dropout = nn.Dropout(dropout)
def split_heads(self, x):
batch, seq_len, _ = x.shape
x = x.view(batch, seq_len, self.num_heads, self.head_dim)
return x.transpose(1, 2)
def combine_heads(self, x):
batch, _, seq_len, _ = x.shape
x = x.transpose(1, 2).contiguous()
return x.view(batch, seq_len, self.num_heads * self.head_dim)
def forward(self, x, causal=True):
q = self.split_heads(self.q_proj(x))
k = self.split_heads(self.k_proj(x))
v = self.split_heads(self.v_proj(x))
scores = q @ k.transpose(-2, -1)
scores = scores / math.sqrt(self.head_dim)
if causal:
seq_len = x.size(1)
mask = torch.tril(
torch.ones(seq_len, seq_len, device=x.device, dtype=torch.bool)
)
scores = scores.masked_fill(~mask, float("-inf"))
weights = torch.softmax(scores, dim=-1)
weights = self.dropout(weights)
output = weights @ v
return self.out_proj(self.combine_heads(output))
The lower-triangular mask prevents position t from reading positions after t. Without it, training leaks the answer into the input.
Build the feed-forward network and decoder block
Each block contains attention plus a position-wise feed-forward network. The feed-forward network usually expands the representation, applies a nonlinear activation, and projects it back:
class FeedForward(nn.Module):
def __init__(self, d_model, expansion=4, dropout=0.0):
super().__init__()
hidden = expansion * d_model
self.net = nn.Sequential(
nn.Linear(d_model, hidden),
nn.GELU(),
nn.Linear(hidden, d_model),
nn.Dropout(dropout),
)
def forward(self, x):
return self.net(x)
class DecoderBlock(nn.Module):
def __init__(self, d_model, num_heads, dropout=0.0):
super().__init__()
self.norm1 = nn.LayerNorm(d_model)
self.attn = MultiHeadSelfAttention(d_model, num_heads, dropout)
self.norm2 = nn.LayerNorm(d_model)
self.ff = FeedForward(d_model, dropout=dropout)
def forward(self, x):
x = x + self.attn(self.norm1(x), causal=True)
x = x + self.ff(self.norm2(x))
return x
This is a pre-normalization layout: normalization occurs before each sublayer. The original paper used a different ordering, often called post-normalization. Both designs exist, so the implementation choice should be explicit.
Rank #4
- System Compatibility Note: 2-slot card, 271x112x39mm, single 8-pin power, 200W TDP. Verify chassis clearance and PSU capacity before purchase.
- Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
- 24GB GDDR6 on 192-Bit Bus: Massive 24GB memory with 456 GB/s bandwidth – ideal for LLMs, AI inference, 3D rendering, and generative design.
- Intel Xe2-HPG Architecture: Built on Intel's next-gen architecture with 20 Xe cores and 160 XMX engines for AI acceleration (197 INT8 TOPS).
- PCIe 5.0 Support: PCI Express 5.0 x16 interface for maximum bandwidth with the latest workstation platforms.
Residual connections give each sublayer a direct path through the network and help optimization. Layer normalization keeps intermediate representations on a more manageable scale.
Assemble the complete model
These are reasonable teaching settings rather than universal defaults:
- Context length: 128 or 256
- Model dimension: 128 or 256
- Layers: 4
- Heads: 4 or 8
- Feed-forward width: four times the model dimension
- Dropout: 0.0 to 0.2
class MiniGPT(nn.Module):
def __init__(self, vocab_size, context_length, d_model=128,
num_heads=4, num_layers=4, dropout=0.1):
super().__init__()
self.context_length = context_length
self.token_embedding = nn.Embedding(vocab_size, d_model)
self.position_embedding = nn.Embedding(context_length, d_model)
self.blocks = nn.ModuleList([
DecoderBlock(d_model, num_heads, dropout)
for _ in range(num_layers)
])
self.norm = nn.LayerNorm(d_model)
self.lm_head = nn.Linear(d_model, vocab_size, bias=False)
# Optional weight tying.
self.lm_head.weight = self.token_embedding.weight
def forward(self, token_ids, targets=None):
batch, seq_len = token_ids.shape
if seq_len > self.context_length:
raise ValueError("Sequence exceeds context length")
positions = torch.arange(seq_len, device=token_ids.device)
x = self.token_embedding(token_ids)
x = x + self.position_embedding(positions)
for block in self.blocks:
x = block(x)
logits = self.lm_head(self.norm(x))
loss = None
if targets is not None:
loss = F.cross_entropy(
logits.reshape(-1, logits.size(-1)),
targets.reshape(-1),
)
return logits, loss
Weight tying is optional. It makes the output projection share parameters with the token embedding. Report parameter counts only after specifying vocabulary size, width, layer count, and whether weights are tied:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsnum_params = sum(p.numel() for p in model.parameters())
print(f"{num_params:,} parameters")
Train the model
from torch.utils.data import DataLoader
device = "cuda" if torch.cuda.is_available() else "cpu"
context_length = 128
train_ds = TextDataset(train_ids, context_length)
val_ds = TextDataset(val_ids, context_length)
train_loader = DataLoader(train_ds, batch_size=32, shuffle=True)
model = MiniGPT(vocab_size, context_length).to(device)
optimizer = torch.optim.AdamW(
model.parameters(), lr=3e-4, weight_decay=0.1
)
for step, (x, y) in enumerate(train_loader):
x, y = x.to(device), y.to(device)
optimizer.zero_grad(set_to_none=True)
logits, loss = model(x, y)
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
optimizer.step()
if step % 100 == 0:
print(f"step={step} loss={loss.item():.4f}")
The learning rate, weight decay, batch size, and clipping value are starting points, not guarantees. Track both training and validation loss, periodically evaluate a fixed prompt, set a random seed for repeatability, and save checkpoints. Do not add mixed precision until the basic implementation is working.
Generate text
@torch.no_grad()
def generate(model, token_ids, max_new_tokens, temperature=1.0, top_k=None):
model.eval()
for _ in range(max_new_tokens):
context = token_ids[:, -model.context_length:]
logits, _ = model(context)
logits = logits[:, -1, :] / temperature
if top_k is not None:
values, _ = torch.topk(logits, min(top_k, logits.size(-1)))
threshold = values[:, [-1]]
logits = torch.where(
logits < threshold,
torch.full_like(logits, float("-inf")),
logits,
)
probabilities = torch.softmax(logits, dim=-1)
next_token = torch.multinomial(probabilities, num_samples=1)
token_ids = torch.cat([token_ids, next_token], dim=1)
return token_ids
Encode a prompt, generate IDs, and decode them back to characters. A temperature below 1 makes output more deterministic; a value above 1 increases randomness. top_k restricts sampling to the most likely tokens. Greedy decoding is useful for debugging but often produces repetitive text.
Test before trusting the output
Check shapes
x = torch.randint(0, vocab_size, (2, 16)).to(device)
logits, loss = model(x, x)
assert logits.shape == (2, 16, vocab_size)
assert loss.ndim == 0
The main transformations should be:
token IDs: (B, T)
embeddings: (B, T, C)
split heads: (B, H, T, D)
scores: (B, H, T, T)
head output: (B, H, T, D)
combined output: (B, T, C)
logits: (B, T, V)
Verify causality
Change a future token and compare the earlier output. Before dropout and at evaluation time, the output at an earlier position should not change. This test is more meaningful than checking only that the mask looks triangular.
Overfit one batch
Train repeatedly on one small batch. The loss should fall sharply. If it does not, inspect the target shift, attention transpose, causal mask, vocabulary size, dtype, and learning rate. A model that cannot overfit a tiny batch is not ready for a full corpus.
Recommended Free Tools
Best Value
- POWERFUL BUSINESS PERFORMANCE – The Dell Precision 3431 is a professional-grade business workstation featuring an Intel Core i5-9500 9th Gen Hexa-Core processor, delivering fast performance, efficient multitasking, and enterprise-level reliability for office environments.
- OPTIMIZED MEMORY & STORAGE FOR PRODUCTIVITY – Equipped with 16GB DDR4 RAM for smooth multitasking and a 1TB SSD, this workstation provides lightning-fast boot times, quick file access, and ample storage for business applications and large datasets.
- PPROFESSIONAL GRAPHICS FOR VISUAL WORKLOADS – Featuring an NVIDIA Quadro P620 2GB graphics card, the Dell Precision 3431 is designed for business professionals, engineers, and creatives who need reliable performance for CAD, 3D modeling, and multi-display setups.
- WINDOWS 11 PRO & ESSENTIAL CONNECTIVITY – Pre-installed with Windows 11 Pro, offering advanced security, remote desktop access, and business-friendly features. Built-in WiFi and Bluetooth ensure seamless connectivity to networks, wireless peripherals, and office devices.
- READY-TO-USE WITH INCLUDED KEYBOARD & MOUSE – Comes with a wired keyboard and mouse, ensuring a plug-and-play setup for immediate productivity in any office or professional workspace.
Check numerical stability
assert not torch.isnan(loss)
assert not torch.isinf(loss)
NaNs commonly indicate an incorrectly broadcast mask, an all-masked attention row, an excessive learning rate, or a mixed-precision problem. Every query position must be able to attend to itself and valid earlier positions.
What the model actually learns
The model minimizes cross-entropy for next-token prediction. It may acquire statistical regularities from the corpus, but a small model trained on a small file will not have broad factual knowledge or reliable reasoning. It is best understood as an inspectable experiment in sequence modeling.
Self-attention does not literally “understand” a sentence. It computes learned interactions between token representations. For an example such as “The dog chased the ball because it was excited,” different heads may assign different weights to earlier tokens, but attention weights alone should not be treated as a complete explanation of model reasoning.
Important trade-offs
From scratch versus built-in components
Manual attention is best for learning tensor shapes and masking. PyTorch's native multi-head attention is better for comparison and later optimization. Production code should generally use optimized primitives rather than educational implementations.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCharacter versus subword tokenization
Character tokenization is transparent but creates longer sequences. Subword tokenization is usually a better practical compromise, at the cost of extra tooling and explanation.
Decoder-only versus encoder-only
Use decoder-only causal attention for generation. Use an encoder-only model for classification or embeddings. Use encoder–decoder architecture for translation, summarization, and other sequence transformations. A Transformer is an architecture family; “GPT” specifically refers to a decoder-only causal language-model approach, while BERT-style systems are encoder-only.
Computational limits
The attention-score matrix has two sequence-length dimensions, so its primary interaction cost grows approximately as O(n²d), where n is sequence length and d is representation width. This is separate from parameter memory, activation memory, optimizer-state memory, and inference-time key/value-cache memory.
A longer context can therefore become a memory bottleneck even when the model has relatively few parameters. Truncating to model.context_length prevents an overflow but discards older context; it is not unlimited context.
What to build next
- Replace character tokenization with a subword tokenizer.
- Add proper validation evaluation and plotted loss curves.
- Add checkpoint saving and resumption.
- Compare your manual attention with PyTorch's implementation.
- Build an encoder-only classifier.
- Use Hugging Face Transformers when you want pretrained models, tokenizers, fine-tuning, or deployment workflows.
- Explore vision or diffusion Transformers, where image patches become tokens.
A curated learning path can also include the original paper, annotated implementations, and compact GPT projects; the build-your-own-AI repository is one such collection.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




