Amid the evolution of Generative AI, the industry narrative is undergoing a profound paradigm shift from "parameters first" to "deployment is king." As the computing power arms race for cloud-based large models intensifies, device giants like Apple and Xiaomi have coincidentally shifted their strategic focus to on-device Small Language Models (SLMs), attempting to reconstruct AI operational logic within the confines of mobile devices. The driving force behind this trend is not mere technical showmanship, but an inevitable choice based on privacy compliance, the economics of inference costs, and millisecond-level interaction experiences. Through a deep analysis of Apple OpenELM's extreme compression architecture and Xiaomi MiLM's engineering breakthroughs in mobile adaptation, it is clear that successfully deploying 3B to 6B parameter models on edge computing nodes relies on aggressive quantization techniques, heterogeneous NPU compute scheduling, and extreme utilization of memory bandwidth. This signifies that smartphones are evolving from simple information terminals into "personal agents" with independent thinking capabilities, able to process complex natural language tasks while completely offline and keeping data on-device. For consumers and developers, understanding this technical path is crucial; on-device AI deployment not only resolves the high marginal costs of cloud inference but also redefines human-computer interaction security boundaries and response speeds through localized data closed-loops, signaling an efficiency revolution in mobile computing where "small" leverages "big."
Core Insights: Why Are Tech Giants Pivoting to "On-Device Small Language Models"?
Over the past 18 months, the narrative logic of the AI industry has undergone a significant turning point: rapidly switching from the "Bigger is Better" era, which pursued infinite expansion of parameter scale, to the "Small is Smart" stage, centered on deployment efficiency. For terminal device giants like Apple and Xiaomi, compressing hundred-billion-parameter models onto mobile phones is not merely a display of technical prowess, but a strategic bet concerning survival.
This shift in focus from cloud-based giant models to Mobile SLM (Small Language Models) is primarily driven by three uncompromising engineering and commercial realities: privacy compliance, inference cost, and interaction latency.
1. Ultimate Privacy Protection (Privacy by Design)
In the application of generative AI, cross-border data transfer and privacy leakage are the biggest compliance risks. For mobile phone manufacturers possessing massive amounts of personal data (SMS, photos, health information), uploading all user requests to the cloud for processing is unacceptable.
On-device SLMs allow data to be processed entirely within a local closed loop. As pointed out in the On-Device Language Models survey, models like Apple's OpenELM and Xiaomi's MiLM are designed to perform inference on the device itself, meaning users' sensitive data does not need to be transmitted over the network, fundamentally avoiding the risk of cloud data leakage. This architecture makes "Personal Agents" possible—AI can read local calendars and emails to provide services without crossing privacy red lines.
2. The Economic Ledger of Inference Cost (Inference Cost)
The inference cost of cloud-based large models is linear: the larger the user base, the higher the GPU computing power consumption. For mobile phone manufacturers with hundreds of millions of active users, if every simple Siri request or Xiao AI command were to call a GPT-4 level model, the continuous inference cost would be astronomical.
By deploying on-device models, manufacturers successfully "transfer" the inference computing power to the user's terminal hardware. Utilizing the phone's built-in NPU (such as Apple's Neural Engine or Qualcomm's Hexagon NPU) for local inference, the manufacturer's marginal cost drops to almost zero. This hybrid computing architecture not only reduces the load on data centers but also makes the large-scale free popularization of AI functions commercially feasible.
3. Zero Latency User Experience (Zero Latency)
In mobile scenarios such as instant messaging and voice assistants, network latency is a killer of user experience. Relying on cloud models means having to go through the entire process of "upload-queue-inference-download," and subject to the network environment, response times are often in the range of seconds or more.
Mobile SLMs eliminate the uncertain variable of the network. Local inference can achieve millisecond-level response speeds, running smoothly even in flight mode or weak network environments. For system-level functions requiring real-time feedback (such as real-time translation and input method prediction), on-device models provide certainty and fluidity that cloud solutions cannot match.
Definition of Technical Scope
It needs to be clarified that the SLM discussed in this article specifically refers to lightweight models optimized for mobile devices, typically with parameter counts between 1B and 7B (such as Apple OpenELM 1.1B or Xiaomi MiLM 6B). This differs from "edge computing" in the broad sense; we focus on how to achieve the implementation of large model capabilities on smartphones constrained by battery power consumption, heat dissipation, and memory bandwidth through model compression and architectural innovation.
Tech Primer: What is SLM and How It Runs on Mobile Devices

Before discussing the strategies of Apple and Xiaomi, we need to clarify core concepts. Small Language Models (SLM) are not merely "stripped-down versions" of large models, but a class of specialized models deeply optimized for on-device computing constraints.
What is SLM? The Evolution of the Definition
In the mobile context, SLM typically refers to generative AI models with parameter counts between 1 billion (1B) and 7 billion (7B) that can be directly deployed on smartphones, laptops, or automotive terminals.
Unlike cloud-based giant models with hundreds of billions of parameters (such as GPT-4 or Claude 3 Opus), SLMs are designed to run under limited computing power (watt-level power consumption) and restricted memory (8GB-16GB RAM) while maintaining high availability. According to industry data, the current deployment strategies of mainstream mobile phone manufacturers are as follows:
- Ultra-lightweight (1B-3B): Such as Apple OpenELM (1.1B/3B), resident in memory, responsible for instant-response system-level tasks (such as text prediction, summarization).
- Standard Level (6B-7B): Such as Xiaomi MiLM (6B), usually loaded on demand when stronger reasoning capabilities are needed.
How Do Phones Run AI? Core Technologies Decoded
To run Transformer architectures on devices like phones where heat dissipation and battery life are extremely sensitive, relying solely on the CPU or GPU is unfeasible. The implementation of SLM mainly relies on two technical pillars: Model Quantization and NPU Acceleration.
1. Aggressive Quantization Technology
Traditional cloud models typically use 16-bit floating-point numbers (FP16) or even 32-bit to store weights, which is a devastating blow to phone memory. SLMs widely adopt 4-bit or even lower quantization technologies.
Apple disclosed in its technical report that to fit a 3B model into an iPhone, they adopted a mixed 2-bit and 4-bit configuration strategy, compressing the average weight to 3.5 bits-per-weight. This means:
- Sharp reduction in memory usage: A 3B model originally required about 6GB of VRAM (FP16), but after 3.5-bit quantization, it only needs about 1.3GB of memory to run.
- Accuracy retention: Through Quantization-Aware Training, the model can still maintain performance close to the original accuracy after significant "slimming down."
2. Heterogeneous Computing and NPU
The NPU (Neural Processing Unit) in mobile SoCs (such as Snapdragon 8 Gen 3 or Apple A17 Pro) is the key to running SLMs. NPUs are designed specifically for matrix multiplication, with energy efficiency far exceeding that of CPUs and GPUs. The inference process of an SLM is compiled into a specific NPU instruction stream, ensuring the phone does not instantly overheat and throttle when generating text.
Cloud LLM vs. On-Device SLM: Core Differences Comparison
To more intuitively understand the application scenarios of the two, here is a comparison across technical dimensions:
Core Metrics | Cloud LLM | On-Device SLM |
|---|---|---|
Typical Parameter Count | > 100B (Hundreds of billions) | 1B - 7B |
Response Latency | High (Affected by network transmission and queuing) | Extremely Low (Millisecond-level response, zero network overhead) |
Privacy & Security | Data must be uploaded to servers | Extremely High (Data never leaves the device) |
Operating Cost | High cost per inference (GPU computing power) | Zero Marginal Cost (Utilizes user's own hardware) |
Network Dependency | Must be online | Fully available offline |
Main Drawbacks | Privacy risks, uncontrollable latency | Limited knowledge base, weaker complex logical reasoning capabilities |
Misconception Correction: Small Models = Stupid?
This is a common misunderstanding. Reducing the number of parameters does indeed compress the world knowledge (Facts) stored in the model, but SLMs are not weak in their execution capabilities for specific tasks.
The core logic of "Small is Smart" lies in data quality. When training Apple's OpenELM and Xiaomi's MiLM, they did not attempt to "memorize the entire internet" like GPT-4, but instead used strictly cleaned textbook-quality data, synthetic data, and domain-specific corpora.
Research shows that through training with high-quality data, small models can rival or even surpass larger models that haven't been optimized for specific tasks in high-frequency scenarios such as summary generation, email polishing, and instruction following. For mobile users, they don't need an AI that can write long sci-fi novels; obtaining an assistant via SLM that perfectly understands the command "send this photo to Mom" is the real experience upgrade.
The Ultimate Showdown: A Deep Comparison of Apple OpenELM and Xiaomi MiLM Architectures
In the On-Device LLM arena, while both Apple and Xiaomi are committed to "fitting models into phones," the technological routes and architectural philosophies behind them present distinctly different approaches. Apple tends to achieve "seamless intelligence" through extreme compression and system-level fusion, whereas Xiaomi focuses more on squeezing out generation capabilities close to cloud-based models within limited computing power.
This section will set aside marketing jargon to conduct a hardcore comparison of Apple OpenELM (and subsequent Foundation Models) and Xiaomi MiLM from three dimensions: model architecture, parameter strategy, and deployment logic.
Apple: The Master of Extreme Compression and Heterogeneous Computing
Apple's strategy is not simply to pursue parameter count, but to pursue the highest intelligence density per unit of energy consumption. Its core architectural design revolves around the Unified Memory Architecture of Apple Silicon, aiming to make the model the underlying driver of iOS system services, rather than just a chatbot.
- Model Specifications and Architecture:
Apple has primarily laid out models at two levels on the mobile end. One is OpenELM, oriented towards open-source research, with a very small parameter count (e.g., 1.1B) and adopting a Layer-wise scaling architecture. According to relevant arXiv research, OpenELM 1.1B improves accuracy by 2.36% compared to similar models while using only half the pre-training tokens.
The second is the ~3B on-device model for actual Apple Intelligence deployment. According to the Apple 2025 Technical Report, this model introduces aggressive technologies such as KV-cache sharing and 2-bit Quantization-Aware Training. This architectural innovation significantly reduces memory usage during inference, allowing it to remain resident in the background without killing foreground applications. - Deployment Ecosystem (MLX):
Apple launched MLX, a machine learning framework optimized specifically for Apple Silicon. It allows developers to fine-tune models like OpenELM directly on local devices. This vertical integration of "hardware-software-model" enables Apple's models to call upon NPU computing power more efficiently when processing token generation, achieving system-level low-latency response.
Xiaomi MiLM: The Balancing Act of Parameter Scale and General Capabilities
Unlike Apple's positioning of "system-level atomic capabilities," Xiaomi's MiLM (Mi Language Model) is more like an "all-around assistant" compressed into a phone. Xiaomi's strategy is to stuff the largest possible model into general mobile platforms like Snapdragon using pruning and quantization technologies to ensure conversation quality.
- Model Specifications and Strategy:
Xiaomi chose two main parameter tiers: 1.3B and 6B. Among them, MiLM-6B is the main model for its flagship phones. According to arXiv survey data, a parameter count of 6B is considered "heavyweight" level for on-device models, giving it an inherent advantage in complex instruction following and long text generation. Xiaomi claims that its self-developed 6B model can rival cloud-based models with tens of billions of parameters in certain scenarios. - Ecosystem Integration (HyperOS):
MiLM is deeply integrated into HyperOS, with its architectural design focusing on cross-device scheduling capabilities. Given Xiaomi's "Human x Car x Home" ecosystem strategy, MiLM must run not only on phone NPUs but also adapt to car chips and IoT devices. Therefore, Xiaomi places greater emphasis on the compatibility of multimodal perception and lightweight deployment in its model architecture, utilizing the heterogeneous computing capabilities of the NPU to balance the power consumption pressure brought by the 6B model.
Architectural Horizontal Comparison
To understand the differences more intuitively, we can compare them across several key technical indicators:
Core Indicator | Apple (OpenELM / On-Device FM) | Xiaomi (MiLM) | Technical Analysis |
|---|---|---|---|
Main Parameter Size | 1.1B / 3B | 1.3B / 6B | Apple pursues extreme lightweighting for memory residency; Xiaomi 6B pursues the upper limit of generation quality. |
Core Architecture | Layer-wise scaling + Grouped Query Attention (GQA) | Standard Transformer variants + Deep Pruning | Apple's architecture leans more towards inference acceleration; Xiaomi's architecture leans towards retaining general capabilities. |
Quantization Tech | 2-bit / 4-bit Mixed Precision | 4-bit Weight Quantization | Apple's 2-bit quantization technology is its key moat for running a 3B model under limited memory. |
Runtime Environment | Core ML / MLX (Highly closed optimization) | Snapdragon NPU / HyperOS (Cross-platform adaptation) | Apple wins on vertical integration efficiency; Xiaomi wins on the breadth of ecosystem hardware. |
In summary, Apple's architectural design is about doing "subtraction," using special architectures (such as KV-cache sharing) to make a 3B model run at the speed of a system component, serving Siri's intent understanding and text polishing; whereas Xiaomi's architectural design is about doing "division," attempting to retain the capabilities of cloud-based large models on the device through quantization compression, serving more complex creation and interaction tasks.
Apple's Strategy: OpenELM Architecture and Ecosystem Closed Loop

Apple's positioning in the SLM (Small Language Model) field demonstrates a typical "hardware-software integration" strategy. Its core is not merely pursuing extreme parameter compression, but achieving efficient on-device inference through the deep coupling of architectural innovation and hardware customization.
1. OpenELM and Layer-wise Scaling Strategy
Unlike cloud-based large models that often reach hundreds of billions of parameters, the OpenELM series models released by Apple in the open-source community (such as Hugging Face) demonstrate its precise calculation of on-device computational boundaries. OpenELM adopts a Layer-wise Scaling strategy, providing four parameter specifications: 270M, 450M, 1.1B, and 3B.
This fine-grained division is not arbitrary but designed to adapt to different computing power levels from Apple Watch to iPhone and then to iPad/Mac:
- 270M/450M: Suitable for ultra-low power background resident tasks, such as text prediction or simple intent classification.
- 1.1B/3B: Mainly targeted at iPhone 15 Pro and subsequent models, handling more complex semantic understanding and generation tasks.
In terms of architectural design, Apple did not adopt a single universal model to solve all problems but instead promoted a "Foundation Model + LoRA Adapter" reference architecture. According to Apple's technical research report, this architecture allows dynamically loading LoRA adapters fine-tuned for specific tasks (such as email summarization, code completion, and style rewriting) on top of a single frozen Foundation Model. This means the device does not need to load a complete model for every function, greatly saving memory usage and improving switching speed.
2. NPU Synergy and Mixed-Precision Quantization
To run 3B-level models under the limited memory bandwidth and battery life of mobile devices, Apple has performed targeted optimizations on the Neural Engine (NPU) of A-series chips (such as A17 Pro).
The most critical technical breakthrough lies in the extreme quantization strategy. Apple developed a mixed-precision scheme that compresses model weights to an average of 3.5 to 3.7 bits while maintaining extremely low accuracy loss. Through the combination of low-bit quantization and LoRA adapters, Apple successfully compressed the model size to a level that can reside in iOS memory without frequent disk swapping (Swap), which is crucial for ensuring Siri's instant response.
3. Privacy First and Ecosystem Closed Loop
Apple's strategic endgame in betting on SLMs lies in Privacy and System-level Integration. By running models like OpenELM on-device, Apple Intelligence ensures that users' personal data (such as messages, calendars, and health records) can be processed without being uploaded to the cloud.
This strategy builds a closed but efficient ecosystem loop:
- Data Security: Only when the on-device SLM determines that a task is too complex (such as generating long texts or complex logical reasoning) will it request the intervention of Private Cloud Compute, and this process is transparent to the user and encrypted.
- System-level Entry Point: SLMs are deeply integrated into iOS's
App Intentsframework, making Siri no longer a simple Q&A bot, but a system-level Agent capable of understanding screen content and operating third-party Apps.
By open-sourcing OpenELM weights and training code, Apple has demonstrated the effectiveness of its architecture to the developer community. This is both a display of technical prowess and a move to attract developers to build more on-device AI applications based on its Core ML framework, further consolidating the moat of its hardware ecosystem.
Xiaomi's Breakthrough: MiLM Model and Hardware-Software Adaptation

Unlike Apple's absolute control over software and hardware, Xiaomi faces more complex fragmentation challenges within the Android ecosystem. To achieve highly available generative AI on compute-constrained mobile devices, Xiaomi adopted a strategy of "lightweight models + deep system integration," the core result of which is the MiLM (Mobile Intelligent Language Model) series.
MiLM Model Matrix and Lightweight Strategy
Xiaomi did not blindly pursue the stacking of parameter scale; instead, addressing the physical limits of mobile memory and power consumption, it launched a model matrix covering different compute levels. According to industry disclosures and related research, Xiaomi has mainly deployed MiLM models in 1.3B and 6B specifications. Among them, the 1.3B version is primarily used for ultra-low power always-on background tasks, such as simple text summarization or system command parsing; while the 6B version (such as MiLM-6B or the subsequent MiMo-V2-Flash) is used to handle complex logical reasoning and creative writing tasks. To fit these models into the limited RAM of mobile phones (typically requiring 4GB-6GB of memory), Xiaomi has carried out aggressive optimization in model quantization, minimizing memory footprint while maintaining precision, thereby establishing its competitiveness in the on-device model field.
HyperOS and Deep NPU Scheduling
At the deployment level, MiLM does not exist as a standalone App but is deeply embedded in the underlying architecture of HyperOS. Due to the diverse chip solutions in the Android camp, Xiaomi must simultaneously adapt to the NPUs (Neural Processing Units) of both the Qualcomm Snapdragon and MediaTek Dimensity platforms. Through heterogeneous computing scheduling, HyperOS can offload MiLM's inference tasks from the CPU/GPU to the NPU, significantly reducing inference latency and heat generation. This hardware-software adaptation allows flagship models to maintain smooth system response speeds when running multi-billion parameter models, avoiding the "lag" and "fast battery drain" issues common in early on-device large models.
XiaoAI and the "Qualitative Change" of the IoT Ecosystem
Xiaomi's unique moat in betting on SLMs lies in its massive AIoT ecosystem. Empowered by MiLM, XiaoAI has evolved from a rule-matching voice assistant into an intelligent Agent with logical understanding capabilities. Users no longer need to memorize fixed command words; the on-device model can understand vague natural language (e.g., "I'm home, make the living room cozy") and automatically translate it into a series of IoT control commands for lights, air conditioners, and speakers. This approach of integrating SLM capabilities down to the device control layer gives Xiaomi stronger practical utility in smart home scenarios compared to pure mobile phone manufacturers.
Performance Verification and Implementation Results
In actual performance tests, MiLM demonstrated deep optimization for the Chinese context. Academic evaluations indicate that Xiaomi MiLM performs robustly in on-device large model benchmarks, especially in Chinese instruction following and context understanding. Compared to cloud-based models, the locally deployed MiLM compromises somewhat on absolute breadth of knowledge, but holds distinct advantages in privacy protection (data does not leave the device) and response speed. This is precisely the value of "AI Pragmatism" that Xiaomi attempts to prove to users on its flagship phones.
Horizontal Review: Comparison Table of Parameters, Performance, and Implementation Scenarios
To provide a more intuitive understanding of the different technical routes taken by Apple and Xiaomi regarding on-device Small Language Models (SLM), we systematically compared their core model parameters, hardware requirements, and main application scenarios. Although the ultimate goal for both is to achieve "intelligence on the phone," in terms of specific implementation paths, Apple has chosen extreme energy efficiency and privacy protection, while Xiaomi tends to provide dialogue capabilities close to the cloud and IoT control on the device.
1. Core Indicators Comparison
The following table summarizes key specification data for Apple's OpenELM series and Xiaomi's MiLM (and subsequent iterations):
Dimension | Apple (OpenELM / Apple Intelligence) | Xiaomi (MiLM / HyperOS AI) |
|---|---|---|
Typical Parameter Scale | 270M, 1.1B, 3B | 1.3B, 6B |
Architecture Features | Layer-wise scaling strategy, deeply optimized for NPU | Lightweight Transformer, adapted for heterogeneous computing on Qualcomm/MediaTek NPUs |
Hardware Threshold | A17 Pro (iPhone 15 Pro) and M-series chips, usually requires 8GB RAM | Snapdragon 8 Gen 3 / Dimensity 9300 and above, 6B models usually require 12GB+ RAM |
Core Advantages | Extreme energy efficiency and privacy: Models are extremely small and can run persistently in the background without seriously affecting battery life. | Strong general capabilities: Larger parameter size; logical reasoning and multi-turn dialogue capabilities are closer to cloud models. |
Main Implementation Scenarios | Siri intent understanding, email summarization, image erasure, code completion (Xcode) | Xiao AI voice interaction, deep document understanding, complex command control for IoT devices |
Ecosystem Openness | Relatively closed (only partial weights like OpenELM are open-sourced on Hugging Face), relies on Core ML/MLX | Relatively open, actively adapts to the Android ecosystem and third-party hardware platforms |
2. Key Differences Analysis: Sufficiency vs. Usability
Through comparison, it can be seen that the choices made by the two manufacturers regarding "parameter magnitude" reflect distinctly different product philosophies:
- Apple's "Invisible" Strategy (Sub-3B Models):
Apple's models are mainly concentrated below 3 billion parameters, and they have even launched a micro model with 270 million parameters. According to a review on arXiv, the 1.1B version of OpenELM improved accuracy by 2.36% while requiring only half the training tokens. This ultra-small parameter strategy is not designed for "conversation," but for "task execution." Apple wants the model to run silently in the background like a system component (such as location services), handling specific tasks like "organizing photo albums" or "rewriting emails," minimizing battery and memory usage to ensure users don't feel lag when using other Apps. - Xiaomi's "All-Rounder" Strategy (6B Models):
Xiaomi has deployed models with up to 6 billion parameters on the device side (such as MiLM-6B). On mobile, 6B is a critical threshold, placing extremely high demands on memory bandwidth and NPU computing power (often requiring phones equipped with large capacity RAM). Xiaomi's aggressive strategy aims to let on-device models undertake more complex tasks originally belonging to the cloud, such as long text reading comprehension or complex logical reasoning. This approach allows "Xiao AI" to possess a higher IQ even in an offline state, capable of handling complex smart home linkage instructions, which aligns with Xiaomi's massive AIoT ecosystem needs.
3. Conclusion: Which One Suits You Better?
- Privacy and Battery Priority Users (Apple Route): If you value data never leaving the phone and want AI features seamlessly integrated into the system (such as Spotlight search, input prediction) without sacrificing battery life, Apple's small parameter strategy offers a better experience. It won't give you a strong feeling of "I am chatting with AI," but you will feel that the phone has become smoother to use.
- Geeks and Heavy Smart Home Users (Xiaomi Route): If you want a phone assistant that truly understands complex natural language instructions, or need to frequently process long document summaries on the phone without uploading to the cloud, Xiaomi's large parameter on-device model can provide a stronger "sense of intelligence." For users with a large number of Mi Home devices, the on-device large model's ability to parse fuzzy instructions will significantly improve the smart home control experience.
Real-world Test: How On-device AI Changes Your Daily Mobile Experience?

As the parameter race returns to the hands of users, the value of SLMs (Small Language Models) is no longer just numbers on benchmark charts, but solving actual pain points through on-device computing power. Unlike cloud-dependent large models, 3B-level models deployed locally on phones (such as Apple's on-device models or Xiaomi's MiLM lightweight version) reconstruct daily experiences primarily across three dimensions: low latency, privacy protection, and offline availability.
Scenario 1: Disconnection and Low Latency—"De-stupidifying" Smart Assistants
In the past, on planes, in underground garages, or subways with weak signals, calling a voice assistant often resulted in "Please check your network connection." The core advantage of on-device SLMs lies in local inference capabilities.
- Offline Execution: Thanks to the local computing power of the NPU, new-generation assistants (such as Siri integrated with Apple Intelligence) can understand complex natural language commands in flight mode, such as "open travel photos from last year" or "turn on power saving mode," without sending speech-to-text requests to the cloud.
- Zero-Latency Response: For system-level operations like setting alarms or opening apps, on-device models eliminate the time for network handshakes and data transmission, achieving millisecond-level interaction responses. ARM's prediction also points out that the real breakthrough for on-device AI lies in context awareness, and this awareness must be immediate and localized.
Scenario 2: Privacy-First Management of Information Overload
Today, with hundreds of messages piling up in the notification center, SLMs act as "local gatekeepers." Unlike cloud-based summaries, on-device processing ensures that sensitive data (such as bank verification codes and private email content) never leaves the device.
- Smart Summaries: The notification summary feature introduced in iOS 18.1 utilizes on-device models to analyze the content of emails or instant messaging apps, extracting core points and displaying them on the lock screen, rather than simply truncating the first two lines of text.
- Privacy Red Lines: For business professionals, this mechanism resolves compliance concerns. For example, the system can read and summarize a Non-Disclosure Agreement (NDA) PDF locally without worrying about the document content being uploaded to third-party servers for training or analysis.
Scenario 3: System-Level Generative Capabilities—More Than Just Chatbots
SLMs have brought generative AI capabilities down to the system level, making it a general-purpose tool rather than a single application.
- Writing Tools: In Mail, Notes, and even third-party apps, users can invoke on-device models to rewrite, proofread, and summarize text. For example, converting a casual spoken draft into a "professional style" email reply with one click, a process completed entirely locally.
- Image and Visual Enhancement: Beyond text, on-device computing power also supports real-time image processing. Apple's Visual Intelligence utilizes the Camera Control button, combining on-device recognition with a cloud knowledge base, to instantly identify restaurant information or event flyers and add them to the calendar. This "understand at a glance" interaction significantly shortens the operation path.
Reality Check: The Gap Between Marketing and Reality
Although the vision is beautiful, the current implementation still faces the physical limitations of the "impossible triangle," and users need to manage their expectations:
- Feature Fragmentation and Regional Restrictions: Current on-device AI is not synchronized globally. For instance, Apple Intelligence is not yet open to the EU and China regions, and initial features are mainly focused on English environments. When domestic users use AI features from Android manufacturers like Xiaomi, they are also often limited by specific models or system versions (such as the adaptation progress of HyperOS).
- Battery Life and Heating: Although SLM parameter counts have been significantly reduced, continuously running 3B-level models remains a huge test for phone cooling and batteries. Device heating and battery drain are still noticeable during high-intensity local image generation or long text processing.
- "Pseudo" All-Rounder: Current on-device intelligence is mainly reflected in the optimization of single tasks (such as photo editing and summarization) and has not yet reached the level of a fully autonomous Agent. Although AI evolving from auxiliary tools to autonomous agents is a clear trend, current mobile AI still mostly requires users to actively trigger and confirm actions, and is still far from "handling all cross-app operations with one sentence."
Challenges and Future: The "Impossible Triangle" of Compute, Heat, and Battery Life

Although Apple and Xiaomi portrayed on-device large models as the core of next-generation smartphones at their launch events, at the engineering implementation level, the laws of physics remain an insurmountable wall. Running models with billions of parameters on a palm-sized, passively cooled mobile device is essentially challenging the "impossible triangle" of computing power, heat generation, and battery life.
Physical Limits: When a 3B Model Meets a Phone Battery
The core contradiction of on-device AI lies in the scarcity of hardware resources. It is less a competition of model intelligence and more a competition of Performance per Watt.
Running a 3B (3 billion parameter) level language model, even after 4-bit quantization, requires occupying about 1.5GB to 2GB of unified memory (RAM). For phones with mainstream configurations of 8GB or 12GB RAM, this means the system must make a difficult choice between "keeping the large model resident" and "killing background apps." Even more severe is the occupation of memory bandwidth—large model inference is a typical memory-intensive task, and frequent parameter reading will instantly spike power consumption.
Heat generation is another intuitive pain point. When high-performance SoCs run large models at full load, power consumption can easily break through 10W, which is a huge thermal load for phones relying on passive cooling. As shown by the heating controversy faced by the iPhone 15 Pro series in its early days, even with an advanced 3nm chip like the A17 Pro, overheating and thermal throttling are inevitable under continuous high load. To meet the energy consumption demands of next-generation AI functions, supply chain news even points out that the iPhone 16 Pro series has to increase battery capacity and improve heat dissipation structures to maintain the experience.
The "Impossible Triangle" of Mobile AI
Under current hardware architectures, on-device large models face three indicators that cannot be satisfied simultaneously:
- Model Size (Intelligence Level): The number of parameters determines the model's understanding and generation capabilities. The more parameters, the smarter the model, but the demand for computing power and memory grows linearly or even exponentially.
- Response Speed (Low Latency): User tolerance for voice assistants or input method suggestions is usually in the millisecond range. Inference speed depends on model parameter size and hardware computing power; larger models inevitably lead to slower token generation speeds, creating a noticeable sense of lag.
- Power Consumption and Heat (Endurance): Pursuing high intelligence and fast response inevitably leads to the SoC running at high frequencies, thereby triggering rapid battery drain and device overheating.
Current solutions mostly compromise between these three. For example, the on-device models currently implemented by Xiaomi and Apple are mostly between 3B-7B parameters, and they extensively use Quantization technology to compress volume, sacrificing some precision in exchange for acceptable power consumption and speed.
Developer Perspective: From CoreML to Hybrid Architecture
For developers, directly invoking on-device computing power is not easy. Although Apple provides low-level acceleration interfaces through CoreML and Qualcomm through AI Stack, balancing model inference with resource preemption by foreground applications remains a challenge. Developers need to utilize the NPU (Neural Processing Unit) to share the pressure on the CPU/GPU, because NPUs are designed for matrix operations and have a much higher energy efficiency ratio than general-purpose processors.
The future mainstream trend is inevitably "Hybrid AI". This is a pragmatic layered processing strategy:
- Edge: Responsible for handling high-frequency, low-compute, and strong privacy-sensitive tasks, such as real-time translation, notification summaries, and photo search. This aligns with the core advantages of Edge AI's low latency and data privacy protection.
- Cloud: When a user issues complex instructions like "write me a detailed travel plan" or "generate a complex 4K poster," the system automatically routes the request to super-large models in the cloud (such as GPT-4 or Claude 3.5 level) for processing.
This architecture not only avoids the physical bottlenecks of mobile hardware but also balances user experience and data security, making it the optimal solution to break the "impossible triangle" at present.







