As large model technology extends from digital text generation to physical interaction, Embodied AI has become the next AI battleground, marking a new watershed for algorithm interviews. For candidates, mastering Transformer architectures or traditional robot kinematics is no longer sufficient; interviewers now prioritize the deep integration of multimodal large models with physical entities. To gain a competitive edge, you must build a composite software-hardware knowledge system: understanding the evolution from "LLM as Planner" modular architectures to end-to-end VLA models like RT-2, mastering Sim-to-Real data gap solutions, and balancing high-frequency control with low-frequency inference at the edge. The true technical barrier lies in utilizing Chain of Thought (CoT) and Affordance constraints to ground large model generalization into precise, safe robotic movements, thereby mitigating physical risks from model hallucinations. Mastering this core knowledge—covering data closed-loops, architecture selection, and deployment constraints—is essential for securing high-paying roles and transforming from a general algorithm engineer into an Embodied AI expert with physical world insight.
Core Landscape: The Knowledge Boundaries of Embodied AI Interviews
Interviews for Embodied AI positions often leave job seekers confused: they are neither like pure large model algorithm roles that only test Transformer and CUDA, nor are they exactly equivalent to traditional robot control roles. To position yourself accurately during an interview, you first need to establish a clear knowledge coordinate system.
Core Definition Cheat Sheet
In the eyes of interviewers, Embodied AI is not simply "robot plus chatbot," but a rigorous systems engineering formula. You can use the following logic to define your tech stack:
Embodied AI = Agent (Brain) + Body (Hardware) + Environment (Interaction)
- Agent (Brain): Responsible for perception, understanding, task planning, and decision-making (e.g., LLM, VLM, VLA).
- Body (Hardware): Actuators and sensors, responsible for converting digital signals into physical actions (e.g., dexterous hands, mobile bases).
- Environment: The feedback loop of the physical world; the core difficulties lie in unpredictability and the Sim-to-Real gap.
The Four Core Quadrants of Interview Assessment
Most technical interview questions for Embodied AI will fall at the intersections of the following four dimensions:
- Model Architecture:
- Core Assessment Point: Whether to adopt a layered architecture (Pipeline), i.e., separation of "Perception-Planning-Control," or an End-to-End architecture, i.e., Vision-Language-Action (VLA) models?
- Essential Knowledge: Understand how models like RT-2, OpenVLA directly map images and language to Action Tokens, and how LLMs exist as Planners in traditional modular solutions.
- Data Strategy:
- Core Assessment Point: Data is currently the biggest bottleneck in Embodied AI. There is a very high probability the interviewer will ask: "How do you solve the problem of scarce real-world data?"
- Essential Knowledge: Familiarity with Sim-to-Real transfer techniques, the data efficiency differences between Imitation Learning and Reinforcement Learning, and how to utilize Synthetic Data for training.
- Perception & Planning:
- Core Assessment Point: Robots must not only "see" objects but also understand their "Affordance" (i.e., how an object can be manipulated).
- Essential Knowledge: 3D visual representation (e.g., applications of NeRF, 3D Gaussian Splatting in robotics), the application of Chain of Thought (CoT) in complex task decomposition, and how to prevent large models from generating "hallucinations" in the physical world (e.g., asking a robot to grasp an object when it has no hands).
- Deployment Constraints:
- Core Assessment Point: The contradiction between the low-frequency inference of large models (e.g., below 10Hz) and the high-frequency requirements of robot control (usually above 500Hz).
- Essential Knowledge: Edge computing, model quantization, inference acceleration, and how to design layered hybrid architectures to decouple high-level decision-making from low-level control.
Comparative Analysis: Traditional Robotics vs. Embodied AI
When answering macro questions like "Why do we need Embodied AI," avoid empty talk like "AI is the future." It is recommended to use the table below for a precise strike from a technical dimension:
Dimension | Traditional Robotics | Embodied AI |
|---|---|---|
Instruction Input | Fixed code instructions or preset trajectories | Open-vocabulary, understanding ambiguous natural language instructions |
Generalization Ability | Limited to specific scenarios and specific objects (Overfitting) | Zero-shot / Few-shot, adapting to unseen objects and environments |
Perception Logic | State estimation, geometric calculation | Semantic understanding, combining commonsense reasoning (e.g., "sponges are soft") |
Core Defects | Lack of flexibility; fails if the environment changes slightly | High inference latency; challenges in precision and stability of action execution |
Typical Algorithms | PID Control, SLAM, Kinematics Solving | Transformer, Diffusion Policy, Imitation Learning |
Master this map, and you master the initiative in the interview—next, we will delve into each quadrant to break down specific interview questions and answering strategies.
Architecture Evolution: From Hierarchical Planning to End-to-End VLA

In Embodied AI interviews, system architecture design is often the first "killer question" thrown out by interviewers. This is not merely an examination of model knowledge, but a test of the candidate's engineering perspective: Do you understand how to effectively combine the semantic understanding capabilities of large models (Brain) with the motion control capabilities of robots (Body)?
Current Embodied AI technical routes are mainly divided into two major schools. The choice between these two architectures directly determines the system's generalization capability, response speed, and training cost:
- Hierarchical Architecture (Modular / Pipeline): This is currently the fastest-to-deploy and most widely applied solution. Its core idea is the separation of the "brain" and the "cerebellum." A Large Language Model (LLM) acts as a high-level Planner, responsible for understanding instructions and decomposing them into a series of sub-tasks; a low-level Controller or Policy network is responsible for executing specific motion controls. This architecture is similar to the human "fast and slow brain" mechanism—the slow brain handles logical reasoning, while the fast brain handles muscle memory.
- End-to-End VLA Architecture (End-to-End Vision-Language-Action): This is a cutting-edge route represented by Google DeepMind RT-2. VLA models attempt to directly map vision and language inputs to low-level robot Action Tokens through a unified Transformer network. This "Pixels-to-Actions" mode breaks the boundary between perception and control, theoretically possessing stronger semantic generalization capabilities.
Interviewers usually test your understanding of technical evolution by comparing these two: from the early SayCan paradigm (LLM + pre-defined skill library) to today's VLA models (multimodal large models directly outputting control signals), every generation of architecture attempts to solve the balance problem between the "semantic gap" and "real-time performance." In the following sections, we will deeply deconstruct the internal mechanisms and interview focus points of these two architectures.
LLM as Planner: How Large Models Act as the "Brain"

In the modular architecture of Embodied AI, the Large Language Model (LLM) plays the role of the "brain," responsible for high-level semantic understanding and task decomposition, while the underlying "cerebellum" or actuators are responsible for specific motion control. This layered architecture is a technical route frequently tested in current interviews; its core lies in how to convert vague natural language instructions into deterministic action sequences executable by robots.
Core Workflow: From Text to API Calls
Interviewers usually ask you to describe a complete reasoning chain. A typical LLM-based Planning workflow includes the following four stages:
- Prompt Engineering: The system encapsulates task instructions (such as "throw away the empty bottle on the table") along with scene descriptions and the robot's capability list (Skill Library) into the System Prompt.
- High-level Planning: The LLM utilizes semantic reasoning capabilities to decompose abstract instructions into an atomic task sequence (Sequence of Primitives). For example:
Find(bottle)->Pick(bottle)->MoveTo(trash_can)->Place(bottle). - Function Calling / API Mapping: The text plan is mapped to specific function calls or API interfaces. This step often combines Code as Policies technology, allowing the LLM to directly output Python code instead of natural language, in order to utilize logic structures like loops and conditional judgments to handle complex tasks.
- Low-level Control: Pre-defined motion primitives or low-level policy networks receive parameters and drive motors to execute physical actions.
Key Technical Points: CoT and ReAct
In this architecture, a simple "Q&A" mode is often insufficient to handle long-horizon tasks; you need to master the following two means of enhanced reasoning:
- Chain-of-Thought (CoT) in Robotics:
In the robotics field, chain-of-thought is not only logical reasoning but also spatio-temporal reasoning. For example, when the instruction is "wash dishes," CoT needs to guide the model to first deduce implied steps like "find sponge" and "apply detergent," rather than directly generating "wash dishes," which is an action that cannot be directly executed. - ReAct (Reasoning + Acting):
This is a high-frequency testing point in interviews. The ReAct mode allows the model to generate a reasoning trajectory (Reason) before generating an action (Act), and observe environmental feedback after execution. For example, the model reasons "I need to pick up the cup," executes the grasping action, and if visual feedback shows "grasp failed," the model will conduct Reflection through the ReAct loop and adjust the strategy, such as "try to readjust the grasping angle." Analysis by Shaqiu Community points out that this dynamic planning capability is key to enhancing Agent autonomy.
Core Challenges: Hallucination and Physical Grounding
The biggest fatal flaw of "LLM as Planner" is Hallucination—that is, the model generates plans that conform to linguistic logic but violate physical constraints (for example, asking a robot to fly to get an item from a high place without a ladder).
To solve this problem, the concept of Affordance must be mentioned in interviews.
- Problem: LLMs only possess semantic knowledge and lack "common sense" of the physical world (such as gravity, object weight, and self-arm span limits).
- Solution: Introduce a Value Function or affordance function as a filter. Taking Google's SayCan as an example, it multiplies the LLM's output ("the probability that I want to do this") with the value function ("the success rate of me doing this now"). An action is executed only when it is both "wanted" and "doable." This mechanism "grounds" high-level planning into physical reality through environmental feedback, preventing the robot from executing dangerous or impossible actions.
Detailed Explanation of VLA Models: Core Principles of RT-2 and PaLM-E

In Embodied AI interviews, interviewers not only focus on whether you understand large models but also on whether you understand how large models interact with the physical world. Vision-Language-Action (VLA) models are cutting-edge architectures that fuse the "brain" (semantic understanding) with the "cerebellum" (motor control). Among them, Google DeepMind's RT-2 and PaLM-E are benchmark cases that must be mastered.
1. Core Mechanism: Action Tokenization
Traditional robot control relies on continuous numerical values (e.g., x=0.25m, y=0.1m), whereas LLMs process discrete tokens. When asked in an interview "how large models directly output control instructions," the core answer lies in action discretization.
- Vocabulary Expansion: VLA models map the robot's action space to text tokens. For example, the range of values for each dimension of a robotic arm's end-effector (x, y, z, roll, pitch, yaw, gripper) is normalized and then divided into fixed intervals (e.g., 256 bins).
- Serialized Output: Actions are no longer simple streams of numerical values but become sequences similar to words.
> Technical Details: In RT-2, a specific action instruction might be represented as a series of tokens, such as<actionx128> <actiony055> <actionz200> ... <terminate>. The model autoregressively predicts these "action words" based on input images and text instructions, just like predicting the next word.
As stated in the People's Daily report on the Lingbao robot, VLA models create an "end-to-end" decision-making system by fusing visual perception, language understanding, and action control, "just like an action version of a large language model."
2. RT-2 Training Strategy: Synergy of Internet Data and Robot Data
The reason RT-2 (Robotic Transformer 2) has become a hot topic in interviews is that it successfully demonstrates the effectiveness of co-training.
- Input Modalities: Images (robot perspective) + Text (task instructions, e.g., "pick up the apple").
- Output Modalities: Text responses (for pure visual question-answering tasks) or Action Tokens (for control tasks).
- Data Ratio:
- Internet-scale Data: Massive amounts of Web text and image data, endowing the model with general semantic understanding capabilities (recognizing what an "apple" is, or who "Taylor Swift" is).
- Robot Trajectory Data: Relatively scarce real-world robot manipulation demonstration data, teaching the model specific physical controls.
- Interview Bonus Point: You can point out that the key to RT-2 lies in retaining the original semantic weights of the VLM (Vision-Language Model), avoiding "catastrophic forgetting" due to fine-tuning on robot data. This allows the model to "transfer" knowledge learned from the internet to physical operations.
3. PaLM-E and "Embodied Chain of Thought"
Unlike RT-2, which directly outputs low-level actions, PaLM-E is viewed more as an embodied multimodal language model.
- Multimodal Sentences: PaLM-E encodes continuous sensor data (such as images, state vectors) into vectors and embeds them directly into the language model's input sequence. The input is not just text, but
Image_Embeddings + "What happened?". - High-level Planning vs. Low-level Control: When distinguishing between the two in an interview, you can explain that PaLM-E is better at generating high-level sequential plans (High-level Plan), such as "go to the kitchen first, then get the cup," which usually requires coordination with a low-level policy to execute specific actions; whereas RT-2 is an end-to-end VLA that directly outputs low-level control signals.
4. Emergent Capabilities
These are the "various cases" that most impress interviewers. Since VLA models have seen various concepts in internet data, they exhibit reasoning capabilities that traditional imitation learning lacks:
- Semantic Reasoning: The instruction is "pick up the extinct animal," and there are a dinosaur toy and a lion toy in the scene. A traditional model would fail if it hasn't seen this instruction, but the VLA model can use semantic knowledge to associate "extinct" with "dinosaur" and execute the grasp.
- Symbol Understanding: The instruction is "put the block on the paper with 'A' written on it." The model can understand the meaning of the character and combine it with spatial location to perform the operation.
Summary Answering Strategy: When answering VLA-related questions, first explain how Tokenization breaks down the barrier between text and action, then use RT-2 as an example to illustrate how semantic knowledge is transferred to the physical world, and finally use examples of emergent capabilities to prove the superiority of this architecture.
Data Closed-loop: Sim-to-Real Transfer and Data Scarcity

In Embodied Intelligence interviews, one of the pain points most frequently examined by interviewers is "Data Starvation." Unlike Large Language Models (LLMs) which possess trillion-level internet text data, the robotics field lacks high-quality, standardized physical interaction data. How to acquire data at low cost, and how to use simulation environments to solve data insufficiency, are key watersheds distinguishing a candidate's engineering implementation ability.
Core Challenges: From "Text Ocean" to "Physical Desert"
LLM's success is built on massive text, but the (State, Action, Reward) data needed by robots is extremely scarce. In interviews, you need to clearly point out this difference:
- High Acquisition Cost: Real-world robot data acquisition relies on Teleoperation or human demonstrations, which involves expensive time costs and safety risks.
- Uneven Distribution: Most existing data is concentrated on simple grasping or moving, lacking data for complex Long-horizon tasks.
Must-Ask Point: Sim-to-Real Transfer Techniques
To solve data scarcity, the industry universally adopts the strategy of "training in simulation, deploying in reality." You need to master the following core techniques and explain how they bridge the "Reality Gap":
- Domain Randomization
This is the most basic and effective method. Its core idea is: if the changes in the simulation environment are drastic enough, the real world is merely a "special case" of the simulation environment.
- Visual Randomization: Randomly change textures, lighting, camera angles, and background noise in simulators (such as Isaac Sim, MuJoCo), forcing the model to learn geometric features of objects rather than visual artifacts.
- Dynamics Randomization: This is a high-level testing point. Not only changing visuals, but also randomizing physical parameters, such as friction coefficients, object mass, joint damping, and motor dead zones. This can greatly improve the robustness of the policy on real hardware.
- System ID & Adaptation
- System Identification: Before deployment, estimate the physical parameters of the real environment through predefined action sequences and align these parameters in the simulation.
- Online Adaptation: Train an extra "Adaptation Module" to adjust the policy network in real-time during inference based on historical observation data (History Window), to cope with motor aging or load changes.
The Data Efficiency Game between Imitation Learning (IL) and Reinforcement Learning (RL)
Interviewers often ask: "In embodied tasks, should one choose Imitation Learning or Reinforcement Learning?" An excellent answer should combine data efficiency and task characteristics:
- Imitation Learning (IL):
- Principle: Based on Behavior Cloning, directly fitting Expert Demonstrations.
- Advantages: Fast training convergence, human-like actions, suitable for initialization of long-horizon tasks.
- Disadvantages: Exists Distribution Shift problem; once errors accumulate, the robot easily enters unseen states and fails. Also, imitation learning is difficult to surpass the level of the demonstrator and lacks the ability to explore new solutions.
- Reinforcement Learning (RL):
- Principle: Maximizing the reward function through Trial and Error.
- Advantages: Possesses exploration capabilities, can discover strategies better than humans, and is more robust to environmental disturbances.
- Disadvantages: Extremely Sample Inefficient, requiring millions of interactions, usually must rely on simulation environments for training before transferring to real machines.
High-Score Strategy: It is recommended to propose an "IL + RL" hybrid paradigm—using Imitation Learning for policy Warm-start to solve the difficulty of initial exploration in RL, and then using RL for Fine-tuning in simulation or on real machines to improve robustness and success rates.
Advanced Topics: Synthetic Data and Data Closed-loop
Besides algorithms, mentioning Data Engineering can reflect your practical experience:
- Synthetic Data: Utilizing generative models (such as Diffusion-based video generation) to expand training data and generate rare Edge Cases, such as object dropping or occlusion scenarios.
- Data Acquisition Hardware: Understanding low-cost data acquisition solutions (such as teleoperation kits composed of mobile phones + robotic arms) is a current industry trend, which lowers the threshold for obtaining high-quality demonstration data from the real world.
When answering such questions, avoid talking only about theoretical formulas. Combine with specific simulators (Isaac Lab, ManiSkill) and troubleshooting experiences of transfer failures (e.g., "The model was perfect in simulation, but oscillated on the real machine due to latency; we solved this problem by domain randomizing latency parameters"), which will significantly increase your credibility.
Engineering Deployment: Inference Latency and Hardware Constraints

In Embodied AI interviews, interviewers not only focus on your understanding of algorithm principles but also value your engineering intuition for deploying large models onto real robots. Candidates with a pure software background often overlook the Real-time and Safety constraints of the physical world, which is the watershed distinguishing a "Paper Reader" from a "Practical Engineer".
Frequency Mismatch: When a 1Hz Brain Meets a 500Hz Body
This is the most classic engineering puzzle in the deployment of Embodied AI. Large Language Models (LLMs) or Vision-Language-Action (VLA) models usually have slow inference speeds, possibly outputting only 1-5 Tokens per second (1-5Hz); whereas a robot's underlying motion control (such as motor servo loops) typically requires a control frequency of at least 50Hz or even 1kHz to ensure smooth and stable movement.
Interviewer often asks: "If your VLA model takes 500 milliseconds to infer once, but the robotic arm needs a control command every 10 milliseconds, how do you fill the Gap?"
Practical Solution: Hierarchical Control
Do not attempt to let the large model drive the motors directly. Mature engineering solutions usually adopt a "Slow Thinking + Fast Execution" hierarchical architecture:
- Upper Layer (Slow Planner): Runs the VLA or LLM, responsible for high-level semantic understanding and task planning (e.g., "go pick up that red cup"), outputting sparse Waypoints or target poses. This layer allows for hundreds of milliseconds of latency.
- Lower Layer (Fast Controller): Runs traditional control algorithms (such as MPC, PID, or Whole-Body Control WBC), responsible for interpolating the targets given by the upper layer into dense motor commands at high frequency (>500Hz).
This architecture not only solves the frequency mismatch but also acts as a safety decoupling, preventing inference lag in the large model from directly causing the robot to "lose control" or jitter.
Compute Power and Latency: The Harsh Reality of Edge Deployment
When running Demos in the laboratory, we often rely on powerful 4090 clusters or cloud APIs, but in industrial sites or on mobile robots, we are often limited by power consumption and heat dissipation, forcing the use of embedded GPUs (like Jetson Orin) or industrial PCs.
Optimization Strategy List:
- Model Quantization: Compress the model from FP32/FP16 to INT8 or even INT4. According to the VLA Model Survey, combining progressive quantization strategies with mixed precision computing (such as FP16/INT8) can reduce calculation volume by 2–4 times and significantly lower memory usage while maintaining performance on benchmark tasks.
- Edge-Cloud Hybrid: Place complex inference that is insensitive to latency (such as long-range task planning, semantic understanding) on the cloud, while deploying models requiring extremely high real-time performance, such as Visual Servoing or obstacle avoidance, on the edge.
- Inference Acceleration Frameworks: Proficient use of TensorRT or ONNX Runtime to perform operator fusion and optimization on models is a basic skill for engineering deployment.
Mini-Case: Edge Optimization of RT-2
Suppose you are asked in an interview: "The RT-2 model has an inference latency of up to 2 seconds on edge devices. How do you optimize it?"
* Analysis: RT-2 is based on ViT and LLM, with a huge number of parameters. A 2-second latency is unacceptable for grasping tasks, as it may lead to grasping empty air after the target object moves.
* Key Answer Points: Besides conventional quantization, you can propose a "Small and Large Model Collaboration" scheme. Use a small, distilled Policy Network (e.g., based on ResNet+MLP) running locally at high frequency to handle specific action execution; only asynchronously call the cloud or backend large RT-2 model for replanning during task switching or when encountering abnormal (OOD) situations.
Safety Constraints: How to Prevent Large Model "Hallucinations" from Causing Injury
In pure NLP tasks, a hallucination might just result in outputting nonsense; but in the robotics field, a hallucination could mean a robotic arm crashing into an operator at full speed.
Essential Engineering Defenses:
- Safety Filter: After the large model outputs action commands, they must pass through a validation layer based on rules or dynamics models. For example, checking if the generated trajectory exceeds joint limits or collides with known obstacles in the environment (Occupancy Map).
- Multimodal Consistency Check: If visual input shows a person ahead, but the language model generates a "full speed ahead" command, the system should possess a safety decoupling mechanism to force an emergency stop or downgrade strategy, rather than blindly trusting the large model.
- Watchdog Mechanism: Targeting network fluctuations or inference timeouts, set up a hardware-level watchdog. Once the heartbeat packet is lost, immediately lock the motor brakes.
In an interview, actively mentioning these "non-AI" traditional robotics assurance mechanisms can greatly demonstrate your respect for deployment risks and your engineering experience.
Selected High-Frequency Interview Questions and Model Answers
In Embodied AI interviews, interviewers not only focus on your memory of algorithm principles but also value how you combine the general capabilities of large models (LLM/VLM) with the specific constraints of robot control. Simply reciting papers makes it hard to pass; you need to demonstrate a profound understanding of the "Perception-Decision-Execution" closed loop.
The following are three of the most representative high-frequency interview questions. It is recommended to answer using the structure of "Core Conclusion + Technical Details + Practical Considerations."
Q1: How did you bridge the Sim-to-Real gap in your projects?
Assessment Point: Examines whether the candidate has practical deployment experience and understands the challenges caused by the differences in dynamics between the Simulator and the physical world.
Reference Answer Logic (Key Talking Points):
- Define the Problem: First, clarify that the core difficulty of Sim-to-Real lies in the simulation environment's inability to perfectly simulate physical friction, lighting changes, and sensor noise in the real world. This leads to models performing excellently in simulation (High Success Rate) but potentially failing completely on real machines.
- Domain Randomization:
- Visual Level: Randomize textures, lighting, and camera angles during training to force the model to learn essential object features rather than environmental backgrounds.
- Dynamics Level: Randomize friction coefficients, object masses, and damping parameters to train a more robust policy.
- Co-fine-tuning and Data Mixing:
- Cite Google DeepMind's research; fine-tuning solely with robot data often lacks generalization. A more effective strategy is Co-fine-tuning with simulation data, real machine data, and internet-scale original web data.
- Emphasize that while real machine data is expensive, it is indispensable and usually used for the final Few-shot Finetuning to align physical characteristics.
- Closed-loop Feedback: Mention introducing Visual or Force Feedback during the inference phase, i.e., correcting actions through real-time observation rather than relying solely on open-loop trajectory prediction.
Q2: Please explain the core differences between RT-1 and RT-2. Why is RT-2 called a VLA model?
Assessment Point: Examines understanding of the evolution of SOTA model architectures, especially the paradigm shift from "traditional Transformer policies" to "Vision-Language-Action (VLA) models."
Reference Answer Logic (Key Talking Points):
- Fundamental Architectural Differences:
- RT-1: Essentially a Transformer-based action generation policy, usually trained from scratch or based on a smaller visual backbone. It focuses mainly on data in the robotics domain, and its generalization ability is limited by the size of the robot dataset.
- RT-2: Is a true VLA (Vision-Language-Action) model. It directly reuses large vision-language models pre-trained on internet-scale data (such as PaLI-X or PaLM-E) as the backbone, treating robot control tasks as a special type of "language generation" task.
- Action Tokenization:
- Explain how RT-2 outputs actions: It discretizes continuous robot actions (such as 6-DoF poses) into Tokens (e.g., quantizing the action space into 256 intervals).
- Key Technical Details: RT-2 extends the VLM's vocabulary by adding dedicated action Tokens. During training, action Tokens and natural language Tokens undergo autoregressive prediction in the same Transformer Decoder. This means the model can utilize pre-trained semantic knowledge (like recognizing a "Superman" toy) to guide unseen manipulation tasks, which is difficult for RT-1 to achieve.
- Inference Constraints:
- Mention the constraint mechanisms during actual deployment: When outputting actions, RT-2 masks out non-action text Tokens to ensure the output consists of valid control instructions, thereby guaranteeing executability.
Q3: How do you evaluate the quality of an Embodied AI model? Is looking at Perplexity enough?
Assessment Point: Examines engineering mindset. Candidates with a background in large models often fall into the NLP evaluation trap, ignoring the physical effectiveness of robotic tasks.
Reference Answer Logic (Key Talking Points):
- Reject Single Metrics: Clearly state that Perplexity or Loss only reflects how well the model fits the training data (accuracy of predicting the next Token), but in embodied scenarios, accurate prediction does not imply the action can be successfully executed.
- Core Business Metrics:
- Success Rate (SR): This is the gold standard. i.e., whether the robot completed the instruction (e.g., "pick up the apple"). This usually requires end-to-end testing on real machines or high-fidelity simulators.
- Sub-goal Completion Rate: For Long-horizon tasks, failure may occur at the very last step. Breaking down sub-goals (e.g., find object -> grasp object -> move -> place) helps locate shortcomings.
- Safety and Efficiency Metrics:
- Executability: Whether the trajectory output by the model complies with kinematic constraints (Kinematics) and whether there are singularity points or collision risks.
- Path Length / Efficiency: The ratio of the trajectory length used to complete the task to the optimal trajectory. A model that spins in circles before finally picking up an object might have an SR of 100%, but its efficiency is extremely low, making it unsuitable for deployment.
- Sim-vs-Real Correlation: In advanced interviews, you can mention that a key focus of evaluation is "whether simulation evaluation results can predict real-machine performance." If the correlation between the two is low, it indicates a failure in building the simulation environment.







