Post-CUDA Era: Under Nvidia's dominance, do the Mojo language and Triton compiler still have a chance?

Jimmy Lauren

Jimmy Lauren

Updated onJan 10, 2026
Read time13 min read

Share

Ace your next interview with real-time, on-screen guidance from GankInterview.

Try GankInterview
Post-CUDA Era: Under Nvidia's dominance, do the Mojo language and Triton compiler still have a chance?

With AI infrastructure firmly controlled by NVIDIA GPUs and the CUDA ecosystem, the concept of a "Post-CUDA Era" does not foretell the hardware giant's demise, but reveals a profound paradigm shift from "hardware-defined software" to "software-defined compute." Amidst the exponential growth of large model parameters and the "Cambrian explosion" of heterogeneous hardware, traditional monolithic development modes can no longer meet complex computing demands. On one hand, AMD, Intel, and various DSA-based NPUs are breaking hardware monopolies, creating an urgent need for a universal, cross-vendor programming solution; on the other, with AI models iterating daily, algorithm engineers cannot tolerate waiting weeks for manual CUDA optimization to verify new operators. This massive gap between compute demand and development efficiency has spawned new programming toolchains represented by Triton and Mojo. This is not merely a change of programming languages, but a revolution in underlying compiler architecture. By leveraging advanced technologies like MLIR, these tools aim to bridge Python's usability and C++'s high performance. Triton, with its pragmatic "block-level programming" model, has significantly lowered the barrier for high-performance operator development, becoming standard in the PyTorch 2.0 era; meanwhile, Mojo ambitiously seeks to unify system-level programming and scripting languages to resolve performance gaps in production environments. The core value of this transformation lies in AI development shifting away from direct reliance on low-level hardware primitives toward a new stage where intelligent compilers drive performance optimization. For developers, understanding this trend is crucial, as future competitive advantage will depend not solely on owning the strongest GPU, but on utilizing advanced software stacks to unleash hardware's ultimate potential within the heterogeneous computing jungle at lower costs and higher efficiency.

What is the "Post-CUDA Era"?

When we talk about the "Post-CUDA Era," this does not mean that NVIDIA's GPUs or the CUDA platform will vanish in the short term. Instead, it is a definition regarding a development paradigm shift: the center of gravity of AI infrastructure is migrating from being "specific hardware (NVIDIA) centric" to "software abstraction (compilers and languages) centric." In this new era, developers are no longer forced to extract hardware performance through tedious low-level CUDA C++ code, but instead achieve cross-hardware efficient computing through higher-level, more general-purpose software stacks.

The formation of this trend is mainly driven by two core forces:

1. The Contradiction Between Hardware Fragmentation and Single-Vendor Lock-in

For a long time, CUDA has built an extremely deep moat, but with the explosion of AI inference and training demands, the hardware market has welcomed a "Cambrian Explosion." Besides NVIDIA, AMD's ROCm ecosystem is maturing rapidly, and Intel as well as various NPUs targeting Domain Specific Architectures (DSA) (such as Huawei Ascend, etc.) are emerging one after another.

In this environment, enterprises and developers can no longer tolerate code being deeply bound to a single piece of hardware. If every new chip requires rewriting the underlying operator library, the R&D cost will become unacceptable. Therefore, the core demand of "Post-CUDA" is portability—that is, "write once, run anywhere," without sacrificing too much performance.

2. The Disconnect Between Model Iteration Speed and Development Barriers

Traditional CUDA programming has an extremely high threshold; developers need to manually manage memory hierarchies, handle Warp synchronization, and deal with complex C++ template metaprogramming. However, the iteration speed of modern AI models is measured in weeks or even days. Algorithm engineers need to quickly verify new operator structures (such as variants of FlashAttention), while waiting for High-Performance Computing (HPC) experts to spend weeks hand-writing optimized CUDA Kernels has become a serious bottleneck.

As industry observers have noted, the real shift is from manual CUDA programming to intelligent compilers generating CUDA-level code. The market urgently needs a toolchain that is as easy to express as Python and as efficient as C++.

In this context, two representative game-changers have emerged:

  • Triton: As a middle-layer abstraction (Intermediate Representation), it focuses on operator development. OpenAI designed it as a DSL based on "Blocks" rather than "Threads," greatly simplifying the difficulty of writing GPU kernels, and it has become a standard feature of PyTorch 2.0.
  • Mojo: As a high-level unified language, it attempts to solve the performance gap problem of Python in production deployment. Mojo aims to provide a system-level solution capable of being downward compatible with hardware details and upward compatible with the Python ecosystem through the MLIR (Multi-Level Intermediate Representation) framework.

The emergence of these two marks that AI development is gradually breaking away from direct reliance on low-level CUDA primitives, entering a new stage where compilers dominate performance optimization.

Triton: The Pragmatist Breaking the Monopoly on Operator Development

Triton: The Pragmatist Breaking the Monopoly on Operator Development

In the exploration of the "post-CUDA era," Triton does not attempt to completely replace CUDA's niche, but instead chooses a more pragmatic path: by raising the abstraction level, it lowers the barrier for Operator development from "hardware experts" to a range accessible by "algorithm engineers."

Triton's core philosophy lies in revolutionizing the traditional GPU programming model. Traditional CUDA development is based on the SIMT (Single Instruction, Multiple Threads) model, where developers must think in terms of "Threads," manually managing synchronization within Warps, the allocation of Shared Memory, and the Coalescing of video memory access. While this fine-grained control can extract maximum hardware performance, it also brings an extremely high cognitive load—one slip-up, and uncoalesced memory access can lead to a precipitous drop in performance.

In contrast, Triton introduces a "Block"-based programming model. Developers no longer get bogged down by the behavior of individual threads but operate directly on data blocks (Tiles). The Triton compiler (built on MLIR) automatically handles complex memory coalescing and shared memory management, "lowering" high-level block operations into efficient machine code. As stated by OpenAI, this design aims to allow researchers without CUDA experience to write extremely efficient GPU code.

The effectiveness of this path has been proven in the industry. PyTorch 2.0 integrates Triton by default as its compiler backend (TorchInductor), marking its transition from an internal research tool at OpenAI to infrastructure within the deep learning ecosystem. Through JIT (Just-In-Time) compilation technology, Triton can dynamically generate kernels at Python runtime, which not only preserves Python's flexibility but also circumvents the tedious process of traditional C++ extension development.

It is worth noting that Triton is not a general-purpose system programming language, but a DSL (Domain Specific Language) focused on tensor computation. It does not attempt to solve all system-level problems but focuses on generating high-performance GPU Kernels. This focus also demonstrates its potential in cross-hardware compatibility. While CUDA is limited to NVIDIA hardware, Triton's backend design allows it to support multiple hardware architectures. Currently, Triton is already capable of running on the AMD ROCm platform, providing developers with a viable path to migrate to non-NVIDIA hardware without rewriting kernel code, which is a crucial step in breaking single-vendor hardware lock-in.

Triton's Core Advantages and Applicable Scenarios

For most AI engineers, writing kernels directly in CUDA C++ often implies extremely high development costs and a steep learning curve. The emergence of Triton is not intended to beat handwritten CUDA in every micro-benchmark, but to find a better balance between Productivity and Performance.

The following is an analysis of Triton's core advantages and best applicable scenarios in actual engineering.

1. Core Advantages: From "Thread Micromanagement" to "Block-Level Abstraction"

Triton's most revolutionary difference lies in its programming model. CUDA requires developers to think from the perspective of a single thread, manually handling warp synchronization, shared memory bank conflicts, and global memory coalescing.

In contrast, Triton adopts a Block-based programming model. Developers operate on data blocks (e.g., tl.load loading a 128×128128 \times 128 matrix block), and the compiler automatically handles underlying instruction parallelization and memory optimization.

  • Automated Memory Optimization: The Triton compiler can automatically analyze data access patterns and optimize memory coalescing, significantly reducing the workload of manually adjusting memory layouts. As pointed out by relevant research, it simplifies the complexity of thread synchronization and vectorization.
  • Native Python Integration: Through the @triton.jit decorator, kernel code is directly embedded in Python projects, supporting Just-In-Time (JIT) compilation without the need to configure complex nvcc build chains.

2. Best Applicable Scenarios

Based on engineering practices and performance evaluations, Triton delivers the most value in the following scenarios:

  • Custom Fused Operators
    This is Triton's "killer" application. In deep learning models, a significant amount of execution time is consumed by VRAM read/write operations (HBM to SRAM) rather than computation itself. By fusing multiple operations (such as LayerNorm + ReLU + Dropout) into a single Triton Kernel, intermediate results can be kept in the on-chip cache (SRAM), significantly reducing bandwidth pressure. PyTorch 2.0's TorchInductor backend defaults to using Triton to generate fused operators precisely based on this logic.
  • IO-Bound Kernel Optimization
    For Softmax, LayerNorm, or simple Element-wise operations, Triton can easily reach peak utilization of hardware bandwidth. Benchmarks show that in scenarios like Fused Softmax, Triton can even beat manually written CUDA code that hasn't been extremely optimized.
  • Attention Mechanism Variants in Large Models (LLMs)
    The popularity of FlashAttention proves the importance of handwritten kernels, but manually maintaining the CUDA version of FlashAttention is extremely painful. Triton enables researchers to quickly implement and test various Attention variants (such as PagedAttention, Sliding Window Attention), with performance often reaching 80%–95% of expert-level CUDA implementations.

3. Complexity and Development Cost Comparison: Triton vs. CUDA

To intuitively demonstrate the differences between the two, we can compare them across two dimensions: code volume and maintenance difficulty:

Dimension

CUDA C++

OpenAI Triton

Programming Paradigm

SIMT (Single Instruction, Multiple Threads)<br>Requires explicit management of threadIdx, blockIdx, shared memory synchronization (__syncthreads()).

SPMD (Single Program, Multiple Data) / Block-wise<br>Operation logic is similar to NumPy; the compiler handles tiling and scheduling.

Code Volume (Vector Add)

~50-100 lines<br>Includes Host/Device memory allocation, Kernel launch configuration, error checking.

~10-15 lines<br>Only need to define computation logic; Python handles data movement.

Performance Ceiling

100% (Theoretical Peak)<br>Can control every register and instruction cycle; suitable for developing foundational libraries like General Matrix Multiplication (GEMM).

~80-95%<br>Slightly inferior on extremely complex compute-intensive operators (like convolution), but often better on fused operators.

Portability

NVIDIA Bound<br>Migration to AMD/Intel requires rewriting in HIP/SYCL.

Gradually Cross-Platform<br>Official support for AMD ROCm exists, and kernels can run without modification.

Summary Recommendation: If you need to squeeze hardware performance to the limit to build general-purpose matrix multiplication libraries (such as a cuBLAS replacement), CUDA remains the undisputed king; but if you are an algorithm engineer aiming to solve specific bottlenecks in models (such as custom Attention or special activation functions), Triton offers a high cost-performance choice of "exchanging 1/5 of the code for 90% of the performance".

Mojo: The Ambitious Unifier of the AI System Stack

Mojo: The Ambitious Unifier of the AI System Stack

In discussions regarding the "Post-CUDA Era," Mojo is often misunderstood as merely another tool for writing GPU Kernels. However, Mojo's positioning is far grander than being a "CUDA alternative." Its core vision is not simply to outperform Nvidia in matrix multiplication, but to solve the long-standing "two-language problem" in AI development—where researchers use Python for prototyping, while system engineers must rewrite code in C++ or CUDA to achieve production-grade deployment.

Mojo attempts to break this barrier by becoming a superset of Python. As stated in the Modular official documentation, Mojo aims to enable developers to achieve system-level performance equivalent to C++ while retaining Python's dynamic features through a progressive approach. This unification not only reduces code migration costs but also allows high-level algorithm engineers to directly access the performance limits of the underlying hardware.

The technical cornerstone supporting this ambition is MLIR (Multi-Level Intermediate Representation). Unlike traditional compiler architectures, MLIR allows Mojo to handle the entire process from high-level graph optimization to low-level hardware instruction generation within the same infrastructure. This enables Mojo to connect "downwards" to various heterogeneous hardware (CPU, GPU, TPU, etc.) while remaining "upwards" compatible with mainstream frameworks like PyTorch.

In terms of its ecosystem niche, Mojo is fundamentally different from Triton. Triton is a Domain-Specific Language (DSL) focused on tensor computation, excelling at writing efficient computation kernels; whereas Mojo aims to handle end-to-end AI pipelines. This includes data preprocessing, model loading, complex control flow logic, and post-processing steps, rather than just matrix operations.

Furthermore, Mojo is not fighting alone; it is part of Modular's commercial landscape. In conjunction with its inference engine MAX (Modular Acceleration X), Mojo acts as a host language to schedule high-performance compute graphs and operators, thereby providing efficiency that surpasses traditional Python runtimes in commercial deployments. In the following sections, we will delve into how Mojo fulfills these performance promises through technical means.

Mojo's "Progressive Typing" and Performance Dividends

Mojo's core technical breakthrough lies in its "Progressive Typing" system, which allows developers to mix dynamic Python code and static system-level code within the same source file. This design is not just for compatibility, but to provide a smooth migration path from prototype to high-performance production environments. To understand how Mojo achieves performance surpassing C++ while maintaining ease of use, we need to delve into its keyword design and compiler architecture.

fn and def: Two Worlds, One Syntax

Mojo introduces the fn keyword to define strictly typed functions, which stands in sharp contrast to the standard Python def and delineates the boundary between dynamic and static:

  • def (Dynamic Mode): Fully compatible with Python behavior. It uses dynamic dispatch, variables can change types at any time, and it is limited by the Global Interpreter Lock (GIL). This is suitable for writing glue code or handling non-compute-intensive logic.
  • fn (Static Mode): This is the source of Mojo's performance. Functions defined by fn are immutable by default, require explicit type declarations, and are directly lowered to machine code via MLIR (Multi-Level Intermediate Representation) at compile time.

In an fn context, the compiler can perform aggressive optimizations because it knows the exact memory layout and types of the data. This means that inside fn, there is no Python interpreter overhead, no GIL interference, and code can run truly in parallel.

Structs and the Ownership Model: Eliminating Overhead

Python's performance bottlenecks stem largely from its object model—almost everything is a heap-allocated PyObject managed via reference counting. Mojo solves this problem by introducing struct and an Ownership model similar to Rust.

  • Value Semantics and struct: Mojo's struct has a static memory layout, typically allocated directly on the stack or inlined in registers. This eliminates the overhead of pointer chasing, allowing data to be arranged compactly in memory, which drastically improves cache hit rates.
  • Zero-copy Mechanism: Through parameter modifiers like borrowed, inout, and owned (ownership transfer), Mojo allows developers to precisely control how data is passed between functions. This enables large tensors or complex data structures to shuttle between functions without memory copying, completely avoiding the "copy-to-modify" performance trap common in Python.

Compiler Optimization: Beyond the "35,000x" Marketing Figure

Although Modular officially demonstrated benchmarks showing Mojo is 35,000 times faster than Python, for systems engineers, it is more important to understand how this acceleration is achieved through the compiler rather than focusing on a single marketing figure. Mojo's performance dividends mainly come from the following underlying compiler optimization capabilities:

  1. Automatic Vectorization (SIMD):
    Because fn provides strict type information, Mojo's compiler (based on LLVM and MLIR) can automatically convert scalar loops into SIMD (Single Instruction, Multiple Data) instructions. Developers can also manually fine-tune using SIMD types in the Mojo standard library to fully utilize modern CPU instruction sets like AVX-512 or AMX.
  2. Tiling and Operator Fusion:
    In AI workloads, memory bandwidth is often the bottleneck. Mojo leverages the multi-level abstraction capabilities of MLIR) to automatically perform Loop Tiling, decomposing large matrix computations into small blocks that fit into CPU/GPU L1/L2 caches. Additionally, it supports Operator Fusion, merging multiple computation steps into a single kernel execution, reducing the frequency of data movement between memory and registers.
  3. Compile-time Metaprogramming:
    Mojo allows parts of the code to be executed during the compilation phase (using alias and comptime). This means tasks like array size checks and specific constant calculations can be completed at compile time. The machine code generated for runtime contains only the core computational logic, thereby achieving zero-overhead abstraction.

Through this "progressive" approach, Mojo effectively encapsulates the complexity of manually optimizing CUDA or C++ code into a Python-like syntax subset, allowing developers to write system-level high-performance code with a lower cognitive burden.

Deep Comparison: Mojo vs. Triton, Competition or Complementary?

Deep Comparison: Mojo vs. Triton, Competition or Complementary?

In discussions about the "post-CUDA era," the most common confusion among developers is: "Should I invest time learning Mojo, or focus on Triton?" This binary perspective often stems from a misunderstanding of their respective niches. In reality, although both Mojo and Triton are challenging NVIDIA's software moat, their abstraction levels, design goals, and positions in the AI stack are distinctly different.

Core Differences: Full-Stack Language vs. Domain Specific Language (DSL)

To understand the relationship between the two, one must first decouple them based on technical attributes. Triton is a DSL embedded in Python, designed to allow non-CUDA experts to write efficient GPU kernels, especially for block-based computations like matrix multiplication and Attention mechanisms. Mojo is a general-purpose system programming language aimed at solving Python's performance defects in deployment and systems programming.

The table below summarizes the key differences between the two:

Feature

Triton

Mojo

Abstraction Level

DSL (Embedded): Highly abstract GPU programming model, enforcing "Block" semantics.

General Language: Superset of Python, supporting full-granularity control from high-level objects to low-level SIMD/registers.

Target Users

Algorithm/Model Engineers: Need custom operators (e.g., FlashAttention) but don't want to touch C++/CUDA complexity.

Systems/Full-Stack Engineers: Building inference engines, compiler infrastructure, or needing to write kernels with complex control flow.

Compilation Mode

JIT (Just-In-Time): Compiles Python AST to PTX/Assembly at runtime, depends on host Python environment.

AOT + JIT: Supports ahead-of-time compilation to standalone binaries, possesses a complete debugger and toolchain, independent of Python runtime.

Applicable Scenarios

Dense Math operators, such as MatMul, Convolution, Attention.

End-to-end AI pipelines: Data loading, preprocessing, model graph execution, and custom kernels.

Niche Analysis: Who is Solving What Problem?

Triton's "sweet spot" lies in operator development efficiency.
For most AI researchers, Triton offers extremely high cost-performance. As shown in relevant benchmarks, although the FlashAttention implemented in Triton might be 10% to 40% slower than hand-written CUDA kernels in peak performance, its code volume is only a fraction of CUDA's, and it is easier to maintain and iterate. Triton automatically handles complex issues like memory coalescing and shared memory management through the compiler, allowing developers to focus on algorithmic logic.

Mojo's ambition lies in unifying the system stack.
Mojo is not just for writing kernels; it is for fixing the fragmentation of the Python ecosystem. In the current architecture, we write models in Python, low-level runtimes in C++, and kernels in CUDA/Triton. Mojo attempts to cover these three levels with one language. According to Modular official documentation, Mojo is a general-purpose language that supports not only AI accelerators but also general programming tasks like CPUs, data transformation, and preprocessing. This means you can write a high-performance HTTP server in Mojo to receive requests, perform data preprocessing, and then invoke model inference, all without expensive context switches between Python and C++.

Complementary Rather Than Substitute: Mojo as Triton's Host

A common misconception is that Mojo will "kill" Triton. In fact, the two are likely to have a symbiotic relationship in the future.

  1. Mojo can serve as Triton's host language: Currently, Triton relies heavily on the Python runtime. With Mojo's compatibility with the Python ecosystem, in the future, we can call Triton kernels directly within Mojo code. This will eliminate the overhead of the Python interpreter (GIL) while retaining Triton's advantages in generating specific operators.
  2. The common language of MLIR: Both Mojo and Triton are based on MLIR (Multi-Level Intermediate Representation) at the bottom layer. Triton lowers Python code to the MLIR triton dialect, and Mojo itself is built on top of MLIR. Theoretically, Mojo's compiler can generate GPU code with performance similar to Triton directly through the MLIR pipeline, or directly integrate Triton's compilation passes.
  3. Handling complex control flow: Triton may be limited by the expressive power of its DSL when handling extremely complex dynamic control flows (such as certain sparse computations or tree search algorithms). In these cases, the low-level control capabilities provided by Mojo (pointer arithmetic, manual memory management, direct register access) can serve as a supplement to Triton, used to handle those "corner cases" that the DSL finds difficult to express.

Conclusion: They are both trying to eliminate the complexity of CUDA, but are exerting force in different dimensions. If your current pain point is "I need to quickly write a variant Attention operator," Triton is the first choice; if you are building the next-generation inference engine, or are troubled by Python's performance bottlenecks in production environments, Mojo is the game-changing tool.

The Common Cornerstone: The Victory of MLIR and Compiler Infrastructure

The Common Cornerstone: The Victory of MLIR and Compiler Infrastructure

When exploring the grand vision of the Mojo language and the immediate capabilities of the Triton compiler, we must look beyond the surface and examine the shared technical backbone supporting both: MLIR (Multi-Level Intermediate Representation). If the CUDA era was a victory of "hand-crafted assembly," then the post-CUDA era is actually a victory of compiler infrastructure.

Why is LLVM Not Enough?

Over the past decade, LLVM has almost dominated compiler backends, but in the face of the explosion of Heterogeneous Computing, traditional LLVM IR has revealed its limitations. LLVM IR is a low-level abstraction that discards high-level semantic information too early—for example, when a matrix multiplication operation is lowered to LLVM IR, it becomes a pile of scalar instructions and memory address calculations. At this level, it is difficult for the compiler to perform complex Loop Tiling, Operator Fusion, or Tensor Layout Optimization, because the original loop structure and data dependencies have become obscured.

The emergence of MLIR is precisely to bridge this gap). It introduces the concept of "multi-level" intermediate representation, allowing compilers to preserve key semantics at different abstraction levels. Through the Dialect mechanism, MLIR allows developers to define domain-specific intermediate representations (such as Linalg for linear algebra, Vector for vectorization, and GPU for heterogeneous acceleration). This means optimization is no longer limited to the quality of the generated machine code but extends back to the algorithmic structure level.

The Art of "Lowering" in Triton and Mojo

Although Triton and Mojo have different positionings, they both rely on MLIR's Progressive Lowering strategy at the bottom layer to achieve cross-hardware high performance and portability.

  • Triton's Path: Triton does not generate PTX directly; instead, it first parses Python-decorated kernels into Triton-IR (an MLIR-based dialect). At this level, the compiler can easily perform Block-level data flow analysis, automatic memory Coalescing, and shared memory conflict elimination. Subsequently, these optimized IRs are progressively lowered to LLVM IR, eventually generating NVIDIA's PTX or AMD's GCN instructions.
  • Mojo's Path: Mojo utilizes MLIR to realize the design of a systems programming language. It can mix high-level Python syntax and low-level system operations within the same source file. The Mojo compiler breaks down high-level structures step-by-step through a series of MLIR Passes, utilizing the Linalg dialect for automatic vectorization and parallelization, while also being able to directly manipulate low-level pointers and registers when necessary.

This architecture makes "write once, run anywhere" no longer an empty slogan. As long as hardware vendors provide corresponding MLIR Dialect backends (such as Intel's SPIR-V or specific NPU backends), the upper-layer code of Triton and Mojo can theoretically be migrated at low cost, without needing to rewrite the entire kernel like in CUDA.

The Paradigm Shift for Developers

For developers on the frontline, understanding this is crucial: Future performance competition is no longer about who can write more ingenious Inline Assembly, but about who can more effectively utilize compiler infrastructure.

In the post-CUDA era, the value of the toolchain lies in automatically generating "Ninja-written" code. As research points out, MLIR-based compilers can achieve over 90% of the performance of hand-written libraries without relying on users to manually specify scheduling strategies or inline assembly. Developers should shift their attention from low-level instruction arrangement to building data flow descriptions that are more suitable for compiler optimization. Whether choosing Triton or Mojo, it is essentially a choice to stand on the shoulders of the giant that is MLIR, utilizing a modular, extensible compiler stack to combat the fragmentation of hardware architectures.

Conclusion: How Should Developers Choose a Tech Stack?

The "Post-CUDA Era" does not mean the end of NVIDIA hardware supremacy, but rather marks a shift in the software development paradigm. NVIDIA GPUs remain irreplaceable in the short term, but CUDA's moat as the sole programming entry point is being dismantled by compiler technologies (like MLIR) and high-level abstract languages (like Triton, Mojo).

For developers at different technical layers, blindly chasing new tools is dangerous. Choosing a tech stack should be based on the project's specific needs for performance limits, development efficiency, and portability. Here are engineering suggestions for different roles:

1. Algorithm Engineers and Operator Optimizers: Master Triton Immediately

If you are a heavy user of PyTorch and your daily work involves large model inference acceleration, FlashAttention variant implementation, or custom operator development, Triton is currently the choice with the highest cost-performance ratio.

  • Applicable Scenarios: Need to bypass PyTorch's heavy scheduling overhead, implement Operator Fusion, or optimize memory access patterns without touching low-level PTX assembly.
  • Return Analysis: According to industry practice, Triton can usually achieve 80% to 95% of the performance of expert-level CUDA kernels with 1/10th of the code volume. Although in extreme optimization scenarios (such as basic primitive library development), hand-written CUDA may still have a 10% to 40% performance advantage, for research projects requiring rapid iteration, Triton's development efficiency is far more important than this marginal performance.
  • Action Suggestion: Do not wait; start rewriting critical bottleneck operators with Triton now.

2. AI Infrastructure Architects and Language Geeks: Keep Watching Mojo

If you focus on the overall architecture of AI systems, model deployment pipelines, or are interested in solving Python's GIL lock and performance bottlenecks, Mojo is a long-term potential stock worth investing time in.

  • Applicable Scenarios: Building unified inference engines, preprocessing pipelines, or attempting to solve the fragmentation problem of "Python prototype + C++ deployment" within a single language.
  • Status Assessment: Mojo has demonstrated exciting system-level programming capabilities, but it is currently still in the early stages, and its ecosystem (especially deep learning framework support) is not as mature as Python's. It is more like a bet on the future, aiming to completely reconstruct the AI software stack through MLIR.
  • Action Suggestion: Stay tuned and try writing some non-critical path high-performance modules (such as data loaders) in Mojo, but avoid fully replacing Python/C++ in production environment core links before its ecosystem is fully mature.

3. Extreme Performance Seekers: Stick to CUDA (But Embrace Generation Tools)

If you are dedicated to developing low-level linear algebra libraries (like cuBLAS level), or your SLA requires squeezing every clock cycle out of the hardware, CUDA remains the only truth. However, even such developers should start paying attention to compiler-assisted tools. Future competition may no longer be about pure hand-written assembly, but how to utilize compiler infrastructure (like TVM or MLIR) to automatically generate high-performance code.

Summary Decision Matrix

Developer Profile

Core Pain Point

Recommended Tech Stack

Reason

PyTorch Algorithm R&D

Existing operators are too slow, memory usage is too high

Triton

Python syntax affinity, high development efficiency, sufficient to cover the vast majority of optimization needs.

System Architect

Python deployment is difficult, C++ development cost is high

Mojo (Watch/Try)

Solves the "Two-Language Problem", unifies training and inference stacks, but needs to wait for ecosystem maturity.

HPC/Library Developer

Pursuing hardware limit performance (P99 Latency)

CUDA + PTX

Fine-grained hardware control (register allocation, instruction pipelining) is still something DSLs cannot completely replace.

Ultimately, the choice of tech stack should not be black and white. For a long time to come, the most robust architecture is likely to be hybrid: writing upper-level logic in Python/Mojo, generating most custom operators with Triton, and retaining hand-written CUDA code only in a very small number of critical paths.

Ace your next interview with real-time, on-screen guidance from GankInterview.

Try GankInterview

Related articles

Stop the prompt superstition: in 2026, the core moat of top Agents is “Harness (control wiring harness)” engineering
Technical Topic•Jimmy Lauren

Stop the prompt superstition: in 2026, the core moat of top Agents is “Harness (control wiring harness)” engineering

If you’re still repeatedly refining prompts for the stability of production-grade AI Agents, the conclusion of this article may overturn you...

Jun 6, 2026
DeepSeek V4 released: a critical first step for open‑source models to “approach GPT.”
Technical Topic•Jimmy Lauren

DeepSeek V4 released: a critical first step for open‑source models to “approach GPT.”

The release of DeepSeek V4 is seen as a key milestone in the history of open-source models because, for the first time, a publicly deployabl...

Apr 27, 2026
DeepSeek V4 Technical Breakdown: What Do MoE + 1M Context Actually Mean?
Technical Topic•Jimmy Lauren

DeepSeek V4 Technical Breakdown: What Do MoE + 1M Context Actually Mean?

DeepSeek V4 introduces a new architecture centered on MoE sparse activation and a 1M context. Its significance for long-sequence reasoning g...

Apr 27, 2026
Behind DeepSeek V4: Chinese AI is taking a different path.
Technical Topic•Jimmy Lauren

Behind DeepSeek V4: Chinese AI is taking a different path.

The emergence of DeepSeek V4 marks China AI’s move onto a path markedly different from mainstream international approaches under constrained...

Apr 26, 2026
Pet System, Internal Codenames, and Employee Emotion Regex: 3 Wild Easter Eggs in Claude Code's Leaked Source Code
Technical Topic•Jimmy Lauren

Pet System, Internal Codenames, and Employee Emotion Regex: 3 Wild Easter Eggs in Claude Code's Leaked Source Code

Recently, the accidental exposure of Anthropic's experimental terminal tool caused an uproar in the developer community. This high-profile C...

Mar 31, 2026
Stop just watching the drama and start learning: From Claude Code's 510,000 leaked lines of code, I learned the state machine architecture of a top-tier Agent.
Technical Topic•Jimmy Lauren

Stop just watching the drama and start learning: From Claude Code's 510,000 leaked lines of code, I learned the state machine architecture of a top-tier Agent.

The recent Claude Code leak is not merely industry gossip, but an invaluable industrial-grade AI engineering blueprint. Deep analysis of the...

Mar 31, 2026