October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Bettesworth Construction
attention

Build Your Own Transformer From Scratch With PyTorch

Build an inspectable decoder-only Transformer in PyTorch. This practical guide covers tokenization, embeddings, causal multi-head attention, decoder blocks, training, testing, and text generation.

By Bettesworth Construction Team 2 min read

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This guide builds a small Transformer neural network from scratch using PyTorch. You will implement attention, multi-head self-attention, positional embeddings, decoder blocks, training, testing, and text generation. It does not cover winding or wiring a mains-voltage electrical transformer. That is a separate electrical-engineering project with serious shock, fire, and insulation hazards.

The finished model is an educational, inspectable decoder-only language model. It can learn patterns from a small text corpus and generate text, but it will not reproduce ChatGPT, Llama, or another production large language model.

What you are building

A Transformer is a neural-network architecture introduced in the 2017 paper “Attention Is All You Need”. Instead of relying primarily on recurrence or convolution to process a sequence, it uses learned, content-dependent interactions between positions.

There are several kinds of Transformer project:

Project Typical use
Attention from scratch Understand queries, keys, values, and attention weights
Encoder-only Transformer Classification, embeddings, and BERT-style tasks
Decoder-only Transformer Autoregressive text generation and GPT-style models
Encoder–decoder Transformer Translation and other sequence-to-sequence tasks
Vision Transformer Image classification using image patches as tokens

This implementation uses the decoder-only design because it provides a clear result: given previous tokens, the model predicts the next token.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
KOTIN Prebuilt Gaming PC RTX 5070 12GB, Ryzen 7 9700X, 32GB DDR5, 1TB SSD
  • POWERED BY RTX 5070 12GB + RYZEN 7 9700X - The GeForce RTX 5070 12GB GDDR7 graphics card pairs with an 8-core AMD Ryzen 7 9700X processor to drive smooth 1440p and 4K gameplay, giving this gaming PC the headroom for modern titles, streaming, and creative work.
  • 32GB DDR5 6000MHz MEMORY & 1TB NVMe SSD - 32GB of high-speed DDR5 memory and a 1TB PCIe 4.0 NVMe solid state drive deliver quick load times, smooth multitasking, and generous storage, keeping this prebuilt gaming desktop responsive under heavy workloads.
  • BUILT-IN 11.3-INCH Smart DISPLAY - An integrated smart screen shows real-time CPU and GPU temperatures, usage, and weather while you play, adding a distinctive and functional touch to your battlestation.
  • 850W 80+ GOLD POWER SUPPLY, 360MM LIQUID COOLING & WiFi 7 - An 850W 80 Plus Gold certified power supply provides stable, efficient power with headroom for future upgrades, while a 360mm AIO liquid cooler, WiFi 7, and an ARGB mid-tower case keep the Ryzen 7 CPU cool and connected in a clean build.
  • READY TO PLAY OUT OF THE BOX - Arrives fully assembled and tested with Windows 11 Home pre-installed, so your prebuilt gaming computer is ready to set up in minutes. Assembled in the USA, and backed by a one-year limited warranty and lifetime free technical support.

How a Transformer represents text

Text must first become numbers. A tokenizer divides text into tokens, then maps each token to an integer ID. Tokens may be characters, words, subwords, or special markers such as beginning-of-sequence, end-of-sequence, and padding tokens.

For the first implementation, character-level tokenization keeps the data pipeline transparent:

text → characters → integer IDs → embeddings → Transformer blocks → logits

Character tokenization has a small vocabulary and requires few dependencies, but it produces longer sequences. Production systems commonly use subword tokenization because word-level vocabularies become very large while character-level sequences can be inefficient.

Install the environment

Use a virtual environment and install the basic dependencies:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
source .venv/bin/activate        # macOS/Linux
.venvScriptsactivate           # Windows PowerShell

python -m pip install --upgrade pip
pip install torch numpy matplotlib tqdm

Check the local installation rather than assuming a particular version:

python --version
python -c "import torch; print(torch.__version__)"
python -c "import torch; print(torch.cuda.is_available())"
python -c "import torch; print(torch.cuda.get_device_name(0) if torch.cuda.is_available() else 'CPU')"

CPU execution is sufficient for shape tests and tiny experiments. Google Colab is another practical option for small educational models. Hardware availability, session duration, CUDA support, and runtime vary, so do not treat one reported GPU workload as a general training estimate.

The architecture

token IDs
   ↓
token embeddings + positional embeddings
   ↓
pre-normalized decoder block
   ├── causal multi-head self-attention
   ├── residual connection
   ├── feed-forward network
   └── residual connection
   ↓
repeat N times
   ↓
final layer normalization
   ↓
linear vocabulary projection
   ↓
next-token logits

For a batch of sequences, use these symbols:

  • B: batch size
  • T: sequence length
  • C: model dimension
  • H: number of attention heads
  • D: head dimension, where D = C / H
  • V: vocabulary size

Prepare a small corpus

Start with a small, legally usable text file. Build a character vocabulary and conversion functions:

Rank #2
YAWYORE Gaming PC Desktop Computer AMD R5 5600GT 16GB 1TB NVMe Towers WiFi
  • Powerful Processor: AMD Ryzen 5 5600GT 3.6GHz (4.6GHz Turbo) 6-Core 12-Thread processor brings faster response time to easily handle multi-threaded tasks
  • Motherboard Specification: MSI A520M-A PRO motherboard provides reliable performance and expandability for your computing needs
  • Integrated Graphics: AMD Radeon Vega Graphics (CPU Integration) enables you to play 1080P mainstream games at quality frame rates
  • Memory and Storage: 16GB DDR4 3200MHz RAM paired with 1TB M.2 NVMe PCIe SSD for fast multitasking and quick data access
  • Power Supply: 550W 80PLUS Bronze certified power supply ensures stable and energy-efficient operation
from pathlib import Path

text = Path("corpus.txt").read_text(encoding="utf-8")
chars = sorted(set(text))
vocab_size = len(chars)

stoi = {ch: i for i, ch in enumerate(chars)}
itos = {i: ch for ch, i in stoi.items()}

encode = lambda s: [stoi[ch] for ch in s]
decode = lambda ids: "".join(itos[i] for i in ids)

ids = encode(text)

Split the underlying text stream before creating overlapping windows. If you create all windows first, nearly identical windows can appear in both training and validation data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Create shifted training examples

A decoder-only language model learns next-token prediction. For:

Input:  The cat sat
Target: cat sat on

the target is shifted one position to the left. The dataset returns a sequence of inputs and the same sequence shifted by one token:

import torch
from torch.utils.data import Dataset

class TextDataset(Dataset):
    def __init__(self, ids, context_length):
        self.ids = ids
        self.context_length = context_length

    def __len__(self):
        return len(self.ids) - self.context_length

    def __getitem__(self, index):
        chunk = self.ids[index:index + self.context_length + 1]
        x = torch.tensor(chunk[:-1], dtype=torch.long)
        y = torch.tensor(chunk[1:], dtype=torch.long)
        return x, y

Every token ID must be an integer tensor with dtype torch.long. The sequence length must not exceed the model's configured context length.

Embeddings and positional information

An embedding table turns each token ID into a learned vector:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
token_embedding = nn.Embedding(vocab_size, d_model)

Self-attention alone does not know order. Without positional information, the same token vectors would represent the same set regardless of sequence order.

The original Transformer used fixed sinusoidal positional encodings. This implementation uses learned positional embeddings because they are simple for a small GPT-style model:

Rank #3
iBUYPOWER Element Gaming PC Desktop Computer Intel Core i7 14700F CPU, NVIDIA GeForce RTX 5070 12GB GPU, 32GB DDR5 RAM, 1TB NVMe SSD, Windows 11 Home, Gamer Keyboard and Mouse - EBI7N5704
  • Intel Core i7 14700F, NVIDIA GeForce RTX 5070 12GB, 32GB DDR5 RGB 4800MHz 16x2 1TB NVMe SSD, WIFI Ready, Windows 11 Home
  • Connectivity: 6 x USB 3.1 | 1x RJ-45 Network Ethernet 10/100/1000 | Audio: On board audio
  • Special Add-Ons: Tempered Glass RGB Gaming Case | 802.11AC Wi-Fi Included | 16 Color RGB Lighting Case | Free iBUYPOWER Gaming Keyboard & RGB Gaming Mouse | No Bloatware | AI Workstation PC ready
self.position_embedding = nn.Embedding(context_length, d_model)

positions = torch.arange(seq_len, device=x.device)
x = token_embedding(token_ids) + self.position_embedding(positions)

Learned positions are tied to the configured maximum context length. Sinusoidal and other positional methods are valid alternatives; neither should be presented as universally best.

Implement scaled dot-product attention

Attention uses queries, keys, and values. A query is compared with keys to produce weights, and those weights form a mixture of value vectors:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Attention(Q,K,V) = softmax((QKᵀ / √dₖ) + M)V

The division by √dₖ matters. As the key dimension grows, unscaled dot products tend to grow in magnitude. Very large softmax inputs produce excessively sharp distributions and weaker useful gradients.

import math
import torch
import torch.nn.functional as F

def scaled_dot_product_attention(q, k, v, mask=None, dropout=None):
    d_k = q.size(-1)
    scores = q @ k.transpose(-2, -1)
    scores = scores / math.sqrt(d_k)

    if mask is not None:
        scores = scores.masked_fill(mask == 0, float("-inf"))

    weights = torch.softmax(scores, dim=-1)

    if dropout is not None:
        weights = dropout(weights)

    return weights @ v, weights

With multiple heads, the expected shapes are:

q:       (B, H, T, D)
k:       (B, H, T, D)
v:       (B, H, T, D)
scores:  (B, H, T, T)
output:  (B, H, T, D)

Add causal multi-head self-attention

Each head attends within a smaller representation subspace. If d_model = 128 and num_heads = 4, each head has dimension 32. The model dimension must divide evenly by the number of heads.

import torch.nn as nn

class MultiHeadSelfAttention(nn.Module):
    def __init__(self, d_model, num_heads, dropout=0.0):
        super().__init__()
        assert d_model % num_heads == 0

        self.num_heads = num_heads
        self.head_dim = d_model // num_heads
        self.q_proj = nn.Linear(d_model, d_model)
        self.k_proj = nn.Linear(d_model, d_model)
        self.v_proj = nn.Linear(d_model, d_model)
        self.out_proj = nn.Linear(d_model, d_model)
        self.dropout = nn.Dropout(dropout)

    def split_heads(self, x):
        batch, seq_len, _ = x.shape
        x = x.view(batch, seq_len, self.num_heads, self.head_dim)
        return x.transpose(1, 2)

    def combine_heads(self, x):
        batch, _, seq_len, _ = x.shape
        x = x.transpose(1, 2).contiguous()
        return x.view(batch, seq_len, self.num_heads * self.head_dim)

    def forward(self, x, causal=True):
        q = self.split_heads(self.q_proj(x))
        k = self.split_heads(self.k_proj(x))
        v = self.split_heads(self.v_proj(x))

        scores = q @ k.transpose(-2, -1)
        scores = scores / math.sqrt(self.head_dim)

        if causal:
            seq_len = x.size(1)
            mask = torch.tril(
                torch.ones(seq_len, seq_len, device=x.device, dtype=torch.bool)
            )
            scores = scores.masked_fill(~mask, float("-inf"))

        weights = torch.softmax(scores, dim=-1)
        weights = self.dropout(weights)
        output = weights @ v
        return self.out_proj(self.combine_heads(output))

The lower-triangular mask prevents position t from reading positions after t. Without it, training leaks the answer into the input.

Build the feed-forward network and decoder block

Each block contains attention plus a position-wise feed-forward network. The feed-forward network usually expands the representation, applies a nonlinear activation, and projects it back:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
class FeedForward(nn.Module):
    def __init__(self, d_model, expansion=4, dropout=0.0):
        super().__init__()
        hidden = expansion * d_model
        self.net = nn.Sequential(
            nn.Linear(d_model, hidden),
            nn.GELU(),
            nn.Linear(hidden, d_model),
            nn.Dropout(dropout),
        )

    def forward(self, x):
        return self.net(x)

class DecoderBlock(nn.Module):
    def __init__(self, d_model, num_heads, dropout=0.0):
        super().__init__()
        self.norm1 = nn.LayerNorm(d_model)
        self.attn = MultiHeadSelfAttention(d_model, num_heads, dropout)
        self.norm2 = nn.LayerNorm(d_model)
        self.ff = FeedForward(d_model, dropout=dropout)

    def forward(self, x):
        x = x + self.attn(self.norm1(x), causal=True)
        x = x + self.ff(self.norm2(x))
        return x

This is a pre-normalization layout: normalization occurs before each sublayer. The original paper used a different ordering, often called post-normalization. Both designs exist, so the implementation choice should be explicit.

Rank #4
ASRock Intel Arc Pro B60 Creator 24GB Graphics Card, Workstation GPU, Xe2-HPG, 2400MHz, 24GB GDDR6 192-bit, PCIe 5.0, 4X DP 2.1, Blower
  • System Compatibility Note: 2-slot card, 271x112x39mm, single 8-pin power, 200W TDP. Verify chassis clearance and PSU capacity before purchase.
  • Dedicated Support: Please contact us directly through Amazon for any product questions or assistance you may require.
  • 24GB GDDR6 on 192-Bit Bus: Massive 24GB memory with 456 GB/s bandwidth – ideal for LLMs, AI inference, 3D rendering, and generative design.
  • Intel Xe2-HPG Architecture: Built on Intel's next-gen architecture with 20 Xe cores and 160 XMX engines for AI acceleration (197 INT8 TOPS).
  • PCIe 5.0 Support: PCI Express 5.0 x16 interface for maximum bandwidth with the latest workstation platforms.

Residual connections give each sublayer a direct path through the network and help optimization. Layer normalization keeps intermediate representations on a more manageable scale.

Assemble the complete model

These are reasonable teaching settings rather than universal defaults:

  • Context length: 128 or 256
  • Model dimension: 128 or 256
  • Layers: 4
  • Heads: 4 or 8
  • Feed-forward width: four times the model dimension
  • Dropout: 0.0 to 0.2
class MiniGPT(nn.Module):
    def __init__(self, vocab_size, context_length, d_model=128,
                 num_heads=4, num_layers=4, dropout=0.1):
        super().__init__()
        self.context_length = context_length
        self.token_embedding = nn.Embedding(vocab_size, d_model)
        self.position_embedding = nn.Embedding(context_length, d_model)
        self.blocks = nn.ModuleList([
            DecoderBlock(d_model, num_heads, dropout)
            for _ in range(num_layers)
        ])
        self.norm = nn.LayerNorm(d_model)
        self.lm_head = nn.Linear(d_model, vocab_size, bias=False)

        # Optional weight tying.
        self.lm_head.weight = self.token_embedding.weight

    def forward(self, token_ids, targets=None):
        batch, seq_len = token_ids.shape
        if seq_len > self.context_length:
            raise ValueError("Sequence exceeds context length")

        positions = torch.arange(seq_len, device=token_ids.device)
        x = self.token_embedding(token_ids)
        x = x + self.position_embedding(positions)

        for block in self.blocks:
            x = block(x)

        logits = self.lm_head(self.norm(x))
        loss = None
        if targets is not None:
            loss = F.cross_entropy(
                logits.reshape(-1, logits.size(-1)),
                targets.reshape(-1),
            )
        return logits, loss

Weight tying is optional. It makes the output projection share parameters with the token embedding. Report parameter counts only after specifying vocabulary size, width, layer count, and whether weights are tied:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
num_params = sum(p.numel() for p in model.parameters())
print(f"{num_params:,} parameters")

Train the model

from torch.utils.data import DataLoader

device = "cuda" if torch.cuda.is_available() else "cpu"
context_length = 128

train_ds = TextDataset(train_ids, context_length)
val_ds = TextDataset(val_ids, context_length)
train_loader = DataLoader(train_ds, batch_size=32, shuffle=True)

model = MiniGPT(vocab_size, context_length).to(device)
optimizer = torch.optim.AdamW(
    model.parameters(), lr=3e-4, weight_decay=0.1
)

for step, (x, y) in enumerate(train_loader):
    x, y = x.to(device), y.to(device)
    optimizer.zero_grad(set_to_none=True)
    logits, loss = model(x, y)
    loss.backward()
    torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
    optimizer.step()

    if step % 100 == 0:
        print(f"step={step} loss={loss.item():.4f}")

The learning rate, weight decay, batch size, and clipping value are starting points, not guarantees. Track both training and validation loss, periodically evaluate a fixed prompt, set a random seed for repeatability, and save checkpoints. Do not add mixed precision until the basic implementation is working.

Generate text

@torch.no_grad()
def generate(model, token_ids, max_new_tokens, temperature=1.0, top_k=None):
    model.eval()
    for _ in range(max_new_tokens):
        context = token_ids[:, -model.context_length:]
        logits, _ = model(context)
        logits = logits[:, -1, :] / temperature

        if top_k is not None:
            values, _ = torch.topk(logits, min(top_k, logits.size(-1)))
            threshold = values[:, [-1]]
            logits = torch.where(
                logits < threshold,
                torch.full_like(logits, float("-inf")),
                logits,
            )

        probabilities = torch.softmax(logits, dim=-1)
        next_token = torch.multinomial(probabilities, num_samples=1)
        token_ids = torch.cat([token_ids, next_token], dim=1)
    return token_ids

Encode a prompt, generate IDs, and decode them back to characters. A temperature below 1 makes output more deterministic; a value above 1 increases randomness. top_k restricts sampling to the most likely tokens. Greedy decoding is useful for debugging but often produces repetitive text.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test before trusting the output

Check shapes

x = torch.randint(0, vocab_size, (2, 16)).to(device)
logits, loss = model(x, x)
assert logits.shape == (2, 16, vocab_size)
assert loss.ndim == 0

The main transformations should be:

token IDs:       (B, T)
embeddings:      (B, T, C)
split heads:     (B, H, T, D)
scores:          (B, H, T, T)
head output:     (B, H, T, D)
combined output: (B, T, C)
logits:          (B, T, V)

Verify causality

Change a future token and compare the earlier output. Before dropout and at evaluation time, the output at an earlier position should not change. This test is more meaningful than checking only that the mask looks triangular.

Overfit one batch

Train repeatedly on one small batch. The loss should fall sharply. If it does not, inspect the target shift, attention transpose, causal mask, vocabulary size, dtype, and learning rate. A model that cannot overfit a tiny batch is not ready for a full corpus.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Dell Precision Workstation PC | Quadro P620 GPU - Editing & Design | Windows 11 Pro | Intel i5-9500 | 16GB RAM 1TB SSD | Home or Office Computer | WiFi 6 AX200 + BT (Renewed)
  • POWERFUL BUSINESS PERFORMANCE – The Dell Precision 3431 is a professional-grade business workstation featuring an Intel Core i5-9500 9th Gen Hexa-Core processor, delivering fast performance, efficient multitasking, and enterprise-level reliability for office environments.
  • OPTIMIZED MEMORY & STORAGE FOR PRODUCTIVITY – Equipped with 16GB DDR4 RAM for smooth multitasking and a 1TB SSD, this workstation provides lightning-fast boot times, quick file access, and ample storage for business applications and large datasets.
  • PPROFESSIONAL GRAPHICS FOR VISUAL WORKLOADS – Featuring an NVIDIA Quadro P620 2GB graphics card, the Dell Precision 3431 is designed for business professionals, engineers, and creatives who need reliable performance for CAD, 3D modeling, and multi-display setups.
  • WINDOWS 11 PRO & ESSENTIAL CONNECTIVITY – Pre-installed with Windows 11 Pro, offering advanced security, remote desktop access, and business-friendly features. Built-in WiFi and Bluetooth ensure seamless connectivity to networks, wireless peripherals, and office devices.
  • READY-TO-USE WITH INCLUDED KEYBOARD & MOUSE – Comes with a wired keyboard and mouse, ensuring a plug-and-play setup for immediate productivity in any office or professional workspace.

Check numerical stability

assert not torch.isnan(loss)
assert not torch.isinf(loss)

NaNs commonly indicate an incorrectly broadcast mask, an all-masked attention row, an excessive learning rate, or a mixed-precision problem. Every query position must be able to attend to itself and valid earlier positions.

What the model actually learns

The model minimizes cross-entropy for next-token prediction. It may acquire statistical regularities from the corpus, but a small model trained on a small file will not have broad factual knowledge or reliable reasoning. It is best understood as an inspectable experiment in sequence modeling.

Self-attention does not literally “understand” a sentence. It computes learned interactions between token representations. For an example such as “The dog chased the ball because it was excited,” different heads may assign different weights to earlier tokens, but attention weights alone should not be treated as a complete explanation of model reasoning.

Important trade-offs

From scratch versus built-in components

Manual attention is best for learning tensor shapes and masking. PyTorch's native multi-head attention is better for comparison and later optimization. Production code should generally use optimized primitives rather than educational implementations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Character versus subword tokenization

Character tokenization is transparent but creates longer sequences. Subword tokenization is usually a better practical compromise, at the cost of extra tooling and explanation.

Decoder-only versus encoder-only

Use decoder-only causal attention for generation. Use an encoder-only model for classification or embeddings. Use encoder–decoder architecture for translation, summarization, and other sequence transformations. A Transformer is an architecture family; “GPT” specifically refers to a decoder-only causal language-model approach, while BERT-style systems are encoder-only.

Computational limits

The attention-score matrix has two sequence-length dimensions, so its primary interaction cost grows approximately as O(n²d), where n is sequence length and d is representation width. This is separate from parameter memory, activation memory, optimizer-state memory, and inference-time key/value-cache memory.

A longer context can therefore become a memory bottleneck even when the model has relatively few parameters. Truncating to model.context_length prevents an overflow but discards older context; it is not unlimited context.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to build next

  1. Replace character tokenization with a subword tokenizer.
  2. Add proper validation evaluation and plotted loss curves.
  3. Add checkpoint saving and resumption.
  4. Compare your manual attention with PyTorch's implementation.
  5. Build an encoder-only classifier.
  6. Use Hugging Face Transformers when you want pretrained models, tokenizers, fine-tuning, or deployment workflows.
  7. Explore vision or diffusion Transformers, where image patches become tokens.

A curated learning path can also include the original paper, annotated implementations, and compact GPT projects; the build-your-own-AI repository is one such collection.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Site Office

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.