How Modern Transformers Work — Part 1

At a high level, modern transformers take an input (text, image, voice) and generate an output (text, image, voice). Simply put: 1 2 3 4 5 6 7 8 tokenizer = Tokenizer() transformer = Transformer() q = "what's your name?" input_idx = tokenizer.encode(q) # [13347, 885, 634, 1308, 30] output_idx = transformer(input_idx) # [15390, 59595, 11, 8113, 656, 166036, 13, 220] o = tokenizer.decode(output_idx) # I'm transformer, built by Amir. Tokenizer We skip explaining the tokenizer piece here, because the goal of this post is the transformer itself. The tokenizer converts raw text into a sequence of ids (using encode()): ...

August 25, 2026 · 9 min · Amir Hadifar

What is a Skill?

A skill is a Markdown file an agent loads on demand to learn how to handle a particular kind of request. It’s useful when you have a repetitive task and don’t want to re-prompt your agent each time. Think of it as the utility function of prompting: instead of duplicating the same instructions in every conversation, you write them once and reuse them. Skills are more general than that, of course — their behaviour adapts to the request in a way a single utility function doesn’t. ...

July 31, 2026 · 5 min · Amir Hadifar

What is MCP (Model Context Protocol)?

This section briefly explains what MCP is and why it’s useful. Before describing MCP, it helps to understand tool-calling first — MCP is built on top of it. Tool-calling Tool-calling is a capability that lets an AI model (like an LLM) interact with the outside world by invoking external functions or APIs. Here’s a simple example of tool-calling in Python with two tools, bash_tool and web_search: 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 # mock tools def bash_tool(command: str) -> str: """Run a shell command and return its output.""" return "index.html main.py styles.css README.md" if command == "ls" else "Done" def web_search(query: str) -> str: """Search the web and return the results.""" return f"Search results for '{query}': Found documentation." tools = {"bash_tool": bash_tool, "web_search": web_search} # Simulates the AI picking a tool based on keywords def mock_llm(query): if "file" in query or "list" in query: return {"tool": "bash_tool", "arguments": {"command": "ls"}} return {"tool": "web_search", "arguments": {"query": query}} while True: user_query = input("User: > ") if user_query.lower() in ["exit", "quit"]: print("Goodbye!") break # Step 1: Get tool choice from LLM decision = mock_llm(user_query) tool_name, args = decision["tool"], decision["arguments"] print(f"AI wants to call: {tool_name}({args})") # Step 2 & 3: execute the tool and print the output tool_output = tools[tool_name](**args) print(f"Tool Output: {tool_output}\n") That’s tool-calling in a nutshell: the model picks a tool (e.g. web_search or bash_tool) and supplies the right arguments (e.g. query: "who won the 2022 World Cup"), the tool or API is executed — locally or remotely — and the model reads back the result. ...

July 20, 2026 · 7 min · Amir Hadifar

Feedforward Neural Networks

1) A brain-inspired analogy As the term “neural network” suggests, these networks are inspired by the computational mechanism of the human brain, where each computational unit is called a neuron. While there’s very little actual resemblance between artificial neural networks and the human brain, the analogy is often used for simplicity’s sake. A biological neuron (top) vs. a neuron in an artificial neural network (bottom) ...

December 9, 2018 · 13 min · Amir Hadifar

Recurrent Neural Networks

When we work with textual data, we’re often dealing with a sequence of characters, words, or sentences, and in most cases the order of that sequence matters to us. Recurrent networks in theory let us represent a sequence of unknown length as a fixed-size vector, while still preserving many of the syntactic and structural properties of the input sequence. Simply, a Recurrent Neural Network (RNN) can be seen as a function that takes an input of length $n$ (e.g., assume a sequence of words) and returns an output vector $y$ of dimension $d_{out}$: ...

September 2, 2018 · 6 min · Amir Hadifar