01/08/2026

Wait... what actually is a programming language?

coding dude

I recently started reading the book "Crafting Interpreters" by Robert Nystrom. And one thing became apparent very quickly - I didn't actually know what a programming language was under the hood. Of course I know what a programming language does and how to use various languages to build and implement things, I don't think I'd be a very good developer if I didn't, what I mean is that very few people actually understand one simple and salient truth: your favourite programming language is simply text in an editor, nothing more. When you write Python code, the computer doesn't 'understand' your Python code, for example. There are many, many complex intermediate steps between the Python code you type into your favourite IDE and your machine actually spitting out any kind of output.

Of course any competent developer at least superficially understands the difference between a compiler and an interpreter, or that "Python is an interpreted language whereas C is a compiled language". But in this article we're going to explore the cases in which those statements fail to hold true, and simply expose a critical lack of understanding as to how computers work.

In order to understand this fully, we're going to use Python as a sort of 'case study'.

If you're reading this, you've probably used Python in some capacity, I picked up Python 2 back in the day as one of my first languages, and used it throughout my time in university for scripting, data science and machine learning. However, up until the last year or so I was pretty much oblivious to Python's different implementations.

Let's say you install Python with a cli package manager, for example homebrew. Which implementation of Python do you get by default? Do you know? For the sake of this post, I'm going to tell you - you get CPython, which is a C implementation of Python. Essentially, you install a binary for your platform which allows your system to interpret Python code and produce some kind of output based on an input you provide, abstracted away semantically and syntactically by what you know as the Python programming language. Again, competent engineers are probably aware that Python has different implementations, Jython, PyPy, etc.. but that's about as far as their knowledge goes. Why would you use Jython, or PyPy, for example, over the default CPython implementation though? Maybe some of you will know when to use them or why you would want to use them, however once you start asking the deeper questions, like how exactly these implementations are built, many of you will start to realise just how much you really don't know about programming languages.

The valley of despair

I started looking into how these implementations are built less than a chapter into Crafting Interpreters, and then closed the book and spent the rest of the week traversing the complex and seemingly infinitely deep rabbit hole that is the subject of programming language implementations. I probably should have just been doing something practical and reading the book, but the deeper I dug, the more interesting I found the subject, so I'm going to share the amalgamation of my research with you in this blog post so that you may spend a day or so wrapping your head around it rather than researching it yourself and spending a week or two.

The first thing that really surprised me is the concept that your Python program, interpreted by the CPython binary, does not produce any native machine code as an output. That's right, nothing you write with Python is actually translated to a machine code instruction and executed by your machine's CPU. It seems ridiculous, or even impossible at first, surely your code must produce a machine code output at some point? And sometimes Python code will be translated to machine code, but not in the case of the CPython implementation^. You see, the CPython implementation already has every code path that can fire as a result of your Python code being interpreted represented as machine code in its binary, you can think of the CPython binary as a translator that takes your Python text and translates it to an already-existing pre-defined code path in the CPython binary.

^ CPython has actually gained its own experimental JIT (which does produce native machine code), and as of 3.14 it ships in the official Windows and macOS builds, but it's disabled by default and still experimental, which is why it doesn't undercut the point above. We will talk more about JIT compilers in the next couple of sections, but you can disregard this for now.

Just to tighten this up so there's no misunderstanding, and this is actually where we throw away the superficial statement of "Python is an interpreted language" too:

The CPython binary is written in C, but your Python code doesn't get translated to C code and executed, what actually happens is an intermediary compilation step - your Python code is compiled to CPython bytecode. This is the CPython binary's first job, take what you have written as Python code (literally plain text in an editor with rules about how it should be written) and compile it to a set of instructions that the CPython interpreter's evaluation loop can then execute. When you hear "interpreters execute code line-by-line" we're talking about the execution of the bytecode as a result of your Python code being compiled into this specific implementation's bytecode. "print("hello world")" for example, is not the line being executed, it is the bytecode representations of that instruction that are being executed line-by-line. This is the thing I want you to nail into your brain at this point - being compiled or interpreted is a property of the language implementation, NOT the language itself.

Saying "Python is an interpreted language" is a fundamentally incorrect statement to make, because the language is not the deciding factor as to whether it is AOT compiled or interpreted, or has a JIT compiler, that's the implementation of the language and the core point you'll need to understand moving forward.

So the full flow - just so we understand how this CPython example works. You write Python in an IDE. The CPython binary then compiles this to CPython bytecode. The interpreter packaged within the CPython binary then interprets the bytecode, line-by-line, through its evaluation loop. The binary (the implementation) is doing all the work. When a line (which is actually an opcode, I'm borrowing the term 'line' to make the idea less abstract) is evaluated via the eval loop, the opcode is fetched from the bytecode, then it is dispatched to its handler, which runs a routine. The evaluation loop then advances to the next opcode.

This sentence: "fetches an opcode, dispatches to its handler, runs it, advances to the next" sounds arcane, but in reality it's much simpler than you think. The eval loop in simplified C is essentially a huge switch statement wrapped in a fetch loop - the core logic is just "fetch instruction", "switch on the instruction" (use opcode as input to switch statement), "jump to the case that you want based on that input", "execute the routine there", and "loop back to fetch the next instruction", until the program is finished. Obviously this is compiled for your specific architecture so this logic lives in the CPython binary. Below is an AI slop diagram that may or may not help you visualise this process.

The slope of enlightenment

You may or may not find my Dunning-Kruger inspired headings amusing, but I do, so that's all that matters. Now we have achieved some semblance of understanding regarding how the CPython implementation above works, we can begin to explore some other popular implementations of the Python programming language. The one I want to focus the rest of this post specifically on is the PyPy implementation. PyPy is a Python implementation written in Python.

Now, you may be thinking, whoaaa didn't you just say earlier that Python code is nothing more than simply text in an editor, with rules?! How the hell can Python be implemented in itself? And you would be right, I said that, and it still holds true, we've just encountered the concept of bootstrapping.

Now, when developers first learn about bootstrapping, it's a fairly straightforward idea. You write a compiler in a different programming language, for your programming language. As a compiler exists for the other programming language, you can compile this compiler with the other language's compiler, and voila - you have a compiler capable of compiling your new language. Once you have this compiler, you can then write a new compiler in the language that you just wrote a compiler for (with the language that had an existing compiler) and boom, your language compiles itself. I think this is a fairly rudimentary CS concept and pretty much all university students understand it, if not - then stay tuned for some of my future blog posts where I will most likely cover the topic in more detail.

The point here, however, is not really to talk solely about bootstrapping compilers, but more the end-to-end flow of how PyPy was achieved, this is something that really helped me to understand the concept of different language implementations thoroughly.

Now, PyPy's name is slightly misleading, as PyPy was not written in the Python that you or I may write in our day-to-day work, but rather a subset of Python called RPython (Restricted Python). RPython is still valid Python code but just uses a smaller subset of the language than you or I may be used to. I'm not going to sit here and pretend that I understand all the nuances between RPython and Python, because quite frankly, I do not. Using the Python example was a means to an end to get my point across in this post, but essentially RPython forbids anything whose meaning isn't fixed until runtime. Why? Because the toolchain has to translate that code to C ahead of time, and you can't statically pin down something that can still change at runtime.

Toolchain? What toolchain? Why does this RPython need to be translated to C ahead of time? Good questions, all will be addressed. To start with, we need to separate our terms, RPython is a language (restricted python). Remember we have a CPython binary, so one would assume that we can use this to execute RPython programs, and one would be correct. Additionally, however, we have something called the RPython toolchain. I'm also not going to describe in detail how the RPython toolchain works in this post, but once I have written a blog post on this I will link it here. The RPython toolchain is a really neat set of tools that allows you to write an interpreter for any dynamic language using RPython. Once you have written the interpreter yourself, and after providing some hints in the form of code annotations, the RPython toolchain works its magic and emits C code. Specifically, the toolchain can emit a full virtual machine and JIT compiler for your language. A C compiler on your system can then compile that to a binary and then you have a language implementation that you have written in RPython compiled to a binary. It's also worth noting that you run the RPython toolchain using the CPython binary (for the first time) in order to, in turn, produce the PyPy binary, as the RPython toolchain itself is written in Python!

Just to really solidify that flow for you:

Climbing the slope

Now, I said I'd explain JIT compilation earlier, and I don't intend on lying to you, however, JIT (just-in-time) compilation could be its own separate blog post so I'll just give you the cliff's notes here. This is the part where the statement regarding compilation vs interpretation being implementation detail rather than being a property of the language itself really starts to make sense. PyPy has a JIT compiler, as we discussed above. This means that hot code paths, i.e., things frequently used over and over in the same way, get compiled down to native machine code at runtime. This is where PyPy really starts to shine for long-running applications. You essentially get the efficiency of a compiled language implementation for certain code hotspots that you simply wouldn't get when using a standard CPython implementation. One caveat on the previous statement is that JIT compilers only really pay off once the program has been running for a long enough time for the tracing JIT to effectively identify hot code paths. This means that for a short-lived script or simple short-running program, you wouldn't really see any kind of performance increase.

For now, this is a good place to stop and have a read yourself, but my next piece will hopefully tie up any loose ends or questions you might have left unanswered from this post. Stay tuned, as I'll be talking lots more about this topic and you can watch as I try and write my own interpreter over the next few posts, peace!