What is the fastest programming language?
Why "what is the fastest programming language?" has no useful answer. We go through the usual ranking of assembly, C, C++, Rust, Java, C# and Python and show what actually decides the speed of a program.

"What is the fastest programming language?" has no useful answer. Speed belongs to a program that runs a given workload on a given machine. The language is only one of many things that decide it.
The usual ranking still comes up in almost every project: assembly first, then C, C++ and Rust, then managed languages like Java and C#, and Python last. Below we go through it and show where each step breaks.
We also discuss this topic in a video:
The usual ranking, and where it breaks
Assembly
Ranked first because you control every instruction and every register. That pays off in a small hot loop, where an expert can still beat the compiler. In a large program, hand-written assembly is rarely worth the cost to write and maintain it.
C
Ranked next because it has almost no runtime. People call C close to the machine, but the compiler rewrites your code heavily. It assumes the program has no undefined behavior and removes code based on that assumption.
C++
Ranked below C because of hidden costs: copies, allocations and virtual calls in hot loops. Written with care, C++ is as fast as C. Exceptions add little overhead on the normal path on most platforms, but throwing one is expensive. The same language gives very different speeds in different hands.
Rust
Rust also uses LLVM for most targets. Like C++, it produces efficient native code, and its place in the ranking depends on the code you write. In safe Rust, the ownership model catches memory errors and data races at compile time. In unsafe code and calls to C through a foreign function interface (FFI), these guarantees no longer hold.
Java
Ranked with C#, for the same reasons: a managed runtime and a GC. The HotSpot JIT compiler uses profile data from the running program to optimize hot code. GraalVM Native Image compiles the program ahead of time, so it starts faster. As with C#, the answer depends on whether you measure startup or a long run.
C#
Ranked with Java because it has a managed runtime and a garbage collector (GC). In practice, the just-in-time (JIT) compiler for x64 and ARM64 is mature. It uses profile data from the running program to optimize hot code. With NativeAOT, the program is compiled to machine code before deployment. It keeps the runtime and the GC, but has no JIT. It starts faster, but long-running code can be a little slower than under the JIT. So even "how fast is C#" has two answers.
Python
Ranked last because the standard interpreter, CPython, runs the code step by step. A loop written in pure Python is slow. But heavy Python code usually calls libraries written in C, C++ or Fortran, such as NumPy. Then the hot loop does not run in Python at all, and the program can be fast.
For every language the honest answer is "it depends". Clang and rustc share one back end, LLVM, but a shared back end does not make two languages equally fast. The code you write and the libraries you use often matter more. Public benchmarks such as the Computer Language Benchmarks Game show the order changing from task to task. If the order changes with the task, it is not a property of the language alone.
Speed without correctness means nothing
A fast program that gives a wrong result is not useful. A speed ranking ignores this, so it compares the wrong thing.
In C and C++ a mistake is easy to make. A wrong pointer, a signed integer overflow or a use after free often compiles without a warning. You pay for the speed with crashes and security bugs. In assembly a mistake is even easier to make and harder to find.
Safe Rust and ordinary C# code prevent most memory corruption. An out-of-bounds index raises an error instead of overwriting memory. Unsafe code and calls into C can still corrupt memory. Integer overflow is still your problem: C# does not check it by default. Safety checks can be cheap when the compiler removes them. It can inline small functions, drop bounds checks it proves safe and turn virtual calls into direct calls.
To remove a bounds check, the compiler must prove that the index stays inside the array. In simple loops it can infer the range of the index and drop the check. Checking the compiler itself is harder. The Alive2 project uses a Satisfiability Modulo Theories (SMT) solver to prove that LLVM optimizations keep the meaning of the program. It has found many real bugs in LLVM this way.
Profile the whole system
In most projects we see, the bottleneck is not the programming language. It is the environment around the program. So before you change the language, profile, and spend the effort where the time actually goes.
Look at the operating system as well as the application. A profiler of your own code will not show system calls, context switches, page faults or time spent waiting for the disk and the network. Linux tools such as perf show them.
The network is a common example. Every packet goes through the kernel network stack, and at high packet rates that cost dominates. Kernel bypass moves the work to user space. DPDK gives the application direct access to the network card, but you write the application against the DPDK API.
The disk is another. If temporary files live on a slow disk, put them on a RAM disk, where they disappear at reboot. If the data must survive, move it to an NVMe SSD. Changes like these often give more than a rewrite in another language.
Remove the work you do not need
Speed has a meaning only for a known workload on known hardware. Much of the overhead in modern programs comes from general-purpose code that pays for cases it never meets. SIMD instructions help, but they are only one step. Remove the work your program does not need:
If the task runs in one thread, do not pay for thread safety. Use single-threaded data structures and a single-threaded allocator.
If threads share no state, do not take locks.
If memory lives for one request, allocate it from one block and free the whole block at the end. Do not free the objects one by one.
If a short-lived tool exits after one job, it does not have to free memory at all. The operating system takes it back.
If the algorithm is fixed and highly parallel, a field-programmable gate array (FPGA) can beat a CPU. It costs more to develop, so check the numbers first. At very high volume, a dedicated chip can pay off.
If a value is known at compile time, compute it at compile time. In C++ that is constexpr. This is the best item on the list: the fastest code is code that never runs.
The one-block rule also explains why a garbage collector costs less than people expect. The .NET GC puts most new small objects in a young area, generation 0. For most small objects, allocation only moves a pointer. Large objects go to a separate large object heap. Most young objects die fast, so the GC frees them in bulk instead of one by one.
The GC still pauses your threads. Young-generation collections stop the program for a short time. Only collections of the oldest generation run mostly in the background, next to your code. Microsoft describes this in its GC fundamentals and background GC pages. Measure the pauses under your real load. If the program has hard real-time deadlines, a GC is usually the wrong choice.
You can take the "do not free" rule to the end and use a GC that never collects. Konrad Kokosa's experimental ZeroGC plugs into the standard .NET runtime. It only moves a pointer and never frees memory. Nethermind's uGC does the same for NativeAOT programs on a single-core RISC-V target, and its core is formally verified with Frama-C. This fits short jobs whose memory fits in RAM: command-line tools, batch jobs, benchmarks. For a long-running service it is the wrong tool. In Kokosa's own test of an ASP.NET Core service, the pauses did not improve, and memory use grew 31 times.
Our view: any language is fine
Using any of these languages is fine. Arguments about which language is the fastest waste time. A professional can make a fast program in almost any of them: they profile, remove the work the program does not need and fit the code to the hardware.
Combining languages works well too. Keep the hot loop in C, C++ or Rust, and write the rest in Python, Java or C#. Python with NumPy already works this way.
There is no fastest programming language
Ask a different question: which language lets us solve this task correctly, on this hardware, within this budget? Then measure on the real target.
If you give up correctness for speed, you will often spend the saved time on bugs.