← All posts

Tagged With

attention

10 posts connected to this tag.

Attention: The Core Of The Transformer

Jun 10, 2026

Attention: The Core Of The Transformer

Attention is the core transformer mechanism: Q/K/V projections, head splitting, RoPE, scaled dot-products, masking, softmax, weighted value sums, GQA, Flash Attention, and the full backward ...

Read post →

Jun 5, 2026

Softmax: The Probability Engine

Lab note Companion post to the Softmax carousel. Previously: Matrix Wx+b: From Scalars To Transformers. The previous post was about the moment a scalar weighted sum grew into a matrix multip...

Read post →

Subscribe

Get my rants delivered to your inbox

I will send new posts as and when I write. No fixed cadence, just engineering notes, rants, and things I am thinking through.

Need an intelligent system to work on real hardware?

Embedded systems · Robotics · Constrained AI · CPU and HPC · Accelerators · Distributed systems

Work With Us / Antshiv Robotics