With the implementation of global data privacy regulations (such as PIPL) and increased awareness of commercial data protection, traditional centralized data aggregation for training is becoming unsustainable. Consequently, securely bridging "data silos" has become a core topic in computational advertising technical interviews. Addressing this challenge requires mastering privacy computing architectures centered on Vertical Federated Learning, enabling joint modeling while strictly adhering to the "data available but not visible" principle. In typical advertising workflows, media platforms holding user behavior features and advertisers possessing high-value conversion labels are often physically isolated; transmitting user IDs or feature tables violates compliance and risks leaking trade secrets. The solution lies in a decoupled distributed training pipeline: First, cryptographic Private Set Intersection (PSI) aligns samples to identify common users without revealing non-intersecting data. Second, during training, Multi-Party Computation (MPC) or Homomorphic Encryption transmits encrypted gradients and intermediate parameters instead of raw data, ensuring participants only access statistics for model updates without reverse-engineering original features. This "data stays, model moves" paradigm secures data sovereignty and privacy while allowing advertisers to leverage high-dimensional media features to optimize Conversion Rate (CVR) prediction models, significantly enhancing precision and commercial returns.
Core Approach: Achieving Joint Modeling via "Data Available but Not Visible"
When answering this question in an interview, you should first present the core principle: "Data available but not visible". Its essence is separating "data ownership" from "data usage rights". Through cryptography and distributed computing technologies, joint model training is completed under the premise that raw data does not leave the local environment (Local Data Retention).
In advertising recommendation scenarios, two parties are usually involved: the Media/Publisher possessing user behavior features and the Advertiser possessing conversion label data. The traditional approach involves physically pooling data from both sides, but this violates current privacy regulations. The privacy computing solution is to keep raw data (such as user ID lists and behavior logs) stored in their respective corporate data silos, exchanging only encrypted intermediate parameters (such as gradients, weights, or ciphertext statistics) over the network.
To achieve this goal, the industry typically adopts a standard technical pipeline. During an interview, this can be summarized into the following three key steps:
- Sample Alignment (Private Set Intersection, PSI):
Before training begins, both parties need to confirm which users are shared (i.e., users who both saw the ad and had subsequent behavior). Utilizing Private Set Intersection (PSI) technology, both parties can safely calculate the intersection of user IDs without revealing their respective non-intersecting user IDs (i.e., unmatched "unique data"), serving as the sample basis for subsequent modeling. - Joint Training (Federated Learning, FL):
Based on the aligned samples, a Vertical Federated Learning (Vertical FL) architecture is adopted. The media party holds features and the bottom model network, while the advertiser holds labels and the top model network. Both parties calculate gradients or intermediate results locally, which are exchanged after encryption or differential privacy processing. As defined in Federated Learning Overview, the core is "data stays, model moves", updating global model parameters through multiple rounds of iteration until the model converges. - Secure Aggregation (Secure Multi-Party Computation, MPC):
During the parameter exchange process, to prevent one party from reverse-engineering the other's raw data, Secure Multi-Party Computation (MPC) or Homomorphic Encryption (HE) is usually combined. For example, utilizing additive homomorphic properties allows parties to see only the aggregated encrypted gradients, rather than specific gradient values from a single participant, thereby mathematically guaranteeing data confidentiality.
This approach not only meets compliance requirements but also breaks down data silos, utilizing multi-party features to improve the accuracy of advertising recommendation models (such as Click-Through Rate CTR or Conversion Rate CVR models).
Technical Breakdown: The End-to-End Privacy Computing Pipeline in Ad Recommendation

In actual interviews for computational advertising, interviewers usually focus on the Vertical Federated Learning architecture. This is because in real business workflows, data often presents a state of "vertical partitioning":
- Media Side: Holds high-dimensional user behavioral features within the App, such as click history, browsing depth, device information, etc.
- Advertiser Side: Holds high-value post-link conversion labels, such as "whether purchased," "whether activated," or "whether deposited."
Both parties hope to utilize each other's data to improve model accuracy (especially CVR prediction), but constrained by data compliance requirements (such as PIPL) and trade secret protection, neither party can directly transmit plaintext User ID lists or wide feature tables.
According to Volcengine's practice in the field of federated learning, the core objective to resolve this contradiction is to build a distributed training pipeline where "data stays, model moves." This end-to-end pipeline usually comprises three decoupled engineering stages: first is privacy alignment (PSI) at the sample layer, ensuring both parties compute based on the same batch of users; second is encrypted gradient exchange at the model layer, where parties only exchange intermediate parameters (such as gradients or residuals) rather than raw features; and finally, collaborative scoring at the inference layer. The following subsections will break down the key technical implementations in this pipeline in detail.
Sample Alignment: Connecting Data Using Private Set Intersection (PSI)

In ad recommendation scenarios within Vertical Federated Learning, the primary engineering challenge faced is not the model architecture, but the problem of Data Silos. Typically, the Media side (Media) holds user behavioral features (such as browsing, clicks, dwell time), while the Advertiser side (Advertiser) holds user conversion labels (such as purchases, registrations).
To train a model predicting Conversion Rate (CVR), these two parts of data must be linked. However, directly exchanging plaintext user IDs (such as phone numbers, IDFA) violates privacy compliance requirements and involves the leakage of trade secrets. At this point, Private Set Intersection (PSI) technology needs to be introduced.
1. Why Not Just Hash Directly?
A common interview question is: "Why not just exchange phone numbers after hashing them with MD5 or SHA256?"
This is a typical security trap. For data with a limited value space like phone numbers or email addresses (e.g., phone numbers only have 11 digits), attackers can easily generate a "Rainbow Table" to reverse-engineer the original plaintext by brute-force matching hash values. Therefore, industrial-grade sample alignment must be based on strong privacy protocols such as asymmetric encryption or Garbled Circuits, ensuring that neither party can learn any non-intersecting data from the other, except for the "intersection users."
2. Core Mechanism and Specific Cases
The core objective of PSI is to calculate the common user set between two parties without leaking non-overlapping data. To avoid confusing interviewers or readers with abstract mathematical formulas, we can understand this through a simple set example:
Scenario Example
* Media Party (Party A) owns User IDs:[User1, User2, User_3]
* Advertiser (Party B) owns Conversion IDs:[User2, User3, User_4]
* Calculation Goal: Both parties identify only[User2, User3]as common samples to enter the training process.
* Privacy Red Line: Party A must never know that Party B possessesUser_4(the advertiser's non-targeted conversion customer); Party B must never know that Party A possessesUser_1(a regular user who did not convert).
3. Technical Implementation: RSA-based Blind Signature Process
In actual engineering implementation, RSA-based Blind Signature or DH (Diffie-Hellman) key exchange are common solutions. Taking the RSA+Hash double-layer encryption mentioned in ByteDance's practice in the field of Federated Learning as an example, the general process is as follows:
- Blinding: The Media Party (Party A) multiplies its User IDs by a random number (blinding) and sends them to the Advertiser (Party B). Due to the existence of the random factor, Party B cannot reverse-engineer the original IDs.
- Signing: Party B uses its own private key to sign this blinded data and returns the result to Party A.
- Unblinding: After receiving the data, Party A uses mathematical properties to remove the previous random number (unblinding). At this point, Party A holds a "list of Party A users signed by Party B's private key."
- Local Calculation and Intersection: Party B also uses the same private key locally to sign the IDs it holds. Finally, both parties exchange (or send one-way) the hashed signed IDs for comparison.
Through this "double-blind" mechanism, neither party can forge the other's IDs, and neither party can restore the original data of the non-intersecting parts through brute-force cracking. This method effectively solves the risk of rainbow table attacks faced by traditional hash alignment and is currently a standard practice in ad attribution and joint modeling.
Model Training: Vertical Federated Learning (Vertical FL) Solves Feature Silos

In advertising recommendation scenarios, the most typical privacy computing requirement is Vertical Federated Learning (VFL). Usually, the media side (Party A) possesses rich user behavior features (Features), while the advertiser or conversion platform (Party B) possesses the final conversion labels (Labels). To train a reasonably effective CTR (Click-Through Rate) or CVR (Conversion Rate) model (such as DeepFM or Wide&Deep), we need to conduct joint modeling without directly exchanging and .
In engineering implementation, this typically adopts a Split Neural Network architecture. We divide the model into two parts: the Bottom Model and the Top Model.
Core Architecture Design
- Bottom Model (Locally Resident): Responsible for processing high-dimensional sparse features. It mainly contains the Embedding layer, mapping sparse features like User ID and Ad ID into dense vectors. These parameters and the original data are completely retained locally by each party and never leave the domain.
- Cut Layer (Interaction Layer): This is the Embedding vector or intermediate layer activation value output by the bottom network. In VFL, only the data of this layer (usually encrypted or obfuscated) will be transmitted.
- Top Model (Server/Label Party): Responsible for feature interaction and loss calculation. It receives Cut Layer inputs from all parties, concatenates them, outputs prediction results through MLP or FM layers, and combines them with labels to calculate Loss.
Detailed Training Process (Round-Trip)
In a typical Mini-batch training iteration, data flow follows the principle of "forward propagation transmits activation values, backward propagation transmits gradients." The following are the standard interaction steps based on encrypted gradients:
- Local Computation
Party A (Feature Party) and Party B (Label Party) first align data locally based on sample IDs. Then, they input feature data into their respective local Bottom Models, perform Embedding lookups and preliminary calculations, obtaining their respective Cut Layer output vectors and . - Encrypted Transmission & Forward
Party A sends its output vector to Party B. To prevent Party B from reverse-engineering A's original features, engineering practices often introduce Homomorphic Encryption or Differential Privacy noise to protect (Note: In some high-performance scenarios, if transmitted via a TEE Trusted Execution Environment, only lightweight obfuscation may be done). After receiving it, Party B concatenates it with its local (Concat) and inputs it into the Top Model to obtain the predicted value . - Global Aggregation & Loss Calculation
Party B holds the label and calculates the Loss (e.g., LogLoss). At this point, the system has completed a full "virtual" forward propagation, just like in a centralized large model, even though the data is physically isolated. - Gradient Return & Update
Entering the backward propagation phase, Party B calculates the gradients of the Top Model and calculates the gradient for Party A's Cut Layer .
- Party B encrypts this gradient and sends it back to Party A.
- Party A decrypts the gradient, combines it with locally cached input data, uses the chain rule to continue backward propagation, and updates the Embedding weights of the local Bottom Model.
- Party B simultaneously updates the weights of its local Top Model and Bottom Model.
Through this mechanism, the two parties only exchange intermediate layer values and gradients. Even the party possessing the labels cannot directly see the other party's original feature data, thereby breaking "feature silos." However, this architecture also brings significant communication overhead. As stated in the 2022 China Privacy Computing Industry Research Report, under homomorphic encryption, the average time for vertical federated learning may be tens or even hundreds of times slower than plaintext training, placing extremely high demands on engineering optimization.
Engineering Practice: How to Solve the Problems of "Slowness" and "Difficulty"
In interviews, being able to derive the gradient formulas for Vertical Federated Learning (VFL) only proves that you understand the algorithmic principles. However, how to deploy this complex interaction process into a production environment is the key to demonstrating the value of a "senior" engineer. Privacy computing faces two core pain points during engineering implementation: "Slow" (performance loss in communication and computation) and "Difficult" (collaboration of complex heterogeneous systems).
First is "Slow". Compared with traditional centralized plaintext training, privacy computing introduces a large amount of cryptographic operations (such as Homomorphic Encryption, MPC) and cross-network communication. According to test data from the 2022 China Privacy Computing Industry Research Report, under the same hardware conditions, the average time required for Vertical Federated Learning modeling can be tens or even hundreds of times slower than plaintext training. In scenarios like ad recommendation that have extremely high requirements for data scale (hundreds of millions of samples) and timeliness (minute-level updates), this performance degradation is unacceptable to business stakeholders.
Second is "Difficult". The complexity of engineering implementation far exceeds that of single-machine model training:
- Offline Training Phase: The bottleneck is usually not in computational power, but in network bandwidth and I/O. Participants need to frequently exchange intermediate gradients or encrypted residuals; network fluctuations on any side will block the entire training link.
- Online Serving Phase: Ad system inference services (Serving) usually have strict SLAs (e.g., < 50ms). Introducing multi-party real-time interaction brings uncontrollable network latency, directly affecting request timeout rates and ad fill rates.
Therefore, solving these problems cannot rely solely on piling up hardware; it requires optimization at the system architecture level. The following content will focus on how to find a viable balance between performance and security through engineering means such as communication compression and asynchronous updates, while ensuring data privacy.
Efficiency Optimization: Communication Compression and Asynchronous Update Strategies
In the engineering implementation of privacy computing, communication overhead is often a more fatal bottleneck than computation overhead. According to test results from the 2022 China Privacy Computing Industry Research Report, in vertical federated learning modeling with 400,000 sample rows × 900 feature columns, the average time consumption can be tens or even hundreds of times slower than plaintext training. To reduce this "unacceptable" latency to a "production-ready" level, deep optimization must be performed on communication bandwidth and synchronization strategies.
1. Communication Compression: Trading Precision for Speed
In advertising recommendation scenarios, model parameters (especially Embedding layers) often exhibit high-dimensional sparse characteristics. Transmitting complete 32-bit floating-point (FP32) gradients not only wastes bandwidth but also brings huge encryption/decryption computational pressure in the encrypted state. The industry typically adopts the following compression strategies:
- Gradient Quantization: Compressing gradients from FP32 to FP16 or even INT8. Although this introduces minor precision loss, in large-scale data training, this noise usually has a negligible impact on the final AUC, yet it can reduce communication volume by 2-4 times.
- Sparsification: For CTR prediction models, not all features are activated in every Batch. By setting a threshold, transmitting only the "Top-k" gradients with large changes, or transmitting only the Embedding gradients corresponding to non-zero features, more than 90% of invalid communication can be filtered out.
- Compressed Sparse Row (CSR) Encoding: Similar to the optimization approach mentioned by Tencent in Best Practices for Recommender Systems, adopting the CSR format to store sparse matrices during the data reading and transmission stages can significantly reduce memory usage and improve GPU utilization.
2. Asynchronous Updates and Pipeline Parallelism
Traditional vertical federated learning mostly adopts a synchronous blocking (Synchronous) mode: Participant A must wait for Participant B to calculate and transmit gradients back before proceeding to the next update. In this mode, the speed of the entire system is limited by the slowest node (Straggler).
- Pipeline Parallelism: To break synchronous waiting, pipeline design is often used in engineering. When the GPU is calculating the gradients of the -th layer (or the -th round), the network I/O thread can simultaneously transmit the ciphertext data of the -th layer (or the -th round). By overlapping computation and communication (Overlap), the latency caused by network transmission is masked.
- Asynchronous SGD: In scenarios with extremely high requirements for timeliness but higher tolerance for model convergence jitter, parties are allowed to use "stale" gradients for updates, no longer strictly waiting for ACKs from all participants. This prevents single-point network fluctuations from blocking the entire training cluster.
3. Hardware-Level Acceleration
In addition to algorithmic optimization, utilizing dedicated hardware to accelerate encryption operations is also a critical path. Since Homomorphic Encryption (HE) involves a large amount of ciphertext matrix multiplication, CPU processing efficiency is extremely low. Currently, top tech giants often introduce FPGA or GPU heterogeneous computing cards to offload these intensive operators. Measured data shows that after introducing FPGA acceleration cards, end-to-end training latency can be reduced by more than 300%, thereby approaching the performance experience of plaintext training as much as possible under the premise of ensuring data privacy (strong security assumptions).
Online Serving: Application of Two-Tower Architecture in Privacy Scenarios

In interviews for ad recommendation systems, interviewers often follow up with a difficult engineering implementation question: "If offline training can use time-consuming MPC (Multi-Party Computation) or Federated Learning, how can we ensure data privacy during the Online Serving stage while meeting the latency requirements of tens of milliseconds for RTB (Real-Time Bidding)?"
1. Conflict Between Real-Time Performance and Strong Privacy Computing
Online ad delivery (RTB) typically requires the entire pipeline to complete within 100ms, covering retrieval, coarse ranking, fine ranking, and bidding logic. However, cryptography-based privacy computing solutions often suffer from significant performance bottlenecks.
- Communication Overhead: Traditional MPC protocols (such as Garbled Circuits or Secret Sharing) require multiple rounds of communication, and the network Round-Trip Time (RTT) will directly blow through the latency budget.
- Computation Overhead: According to the 2022 China Privacy Computing Industry Research Report, in scenarios involving Homomorphic Encryption or Vertical Federated Learning, computation time can be tens or even hundreds of times longer than plaintext computation.
Directly applying heavy privacy protocols to online inference is almost unfeasible under current hardware conditions. Therefore, the mainstream solution in the industry is architectural isolation and hardware acceleration, rather than relying solely on software-level cryptographic protocols.
2. Two-Tower Architecture Under Privacy Protection
To solve performance issues, engineering typically adopts the classic Two-Tower Model (Dual-Encoder), decomposing the computation into "User Side" and "Ad Side" parts, leveraging the computational advantage of vector dot products.
- User Tower: Deployed on the client side (on-device inference) or in a Trusted Proxy.
- Input: Sensitive plaintext data such as the user's real-time behavior sequence and geolocation.
- Output: A low-dimensional dense vector (User Embedding).
- Privacy Isolation: Raw feature data never leaves the user side or the trusted domain; only the Embedding vector, which has undergone non-linear transformation, is output. Although Embeddings theoretically carry the risk of reverse engineering, in engineering practice, they are usually combined with Differential Privacy (DP) to add noise to the vector, or restricted via protocols to be transmitted only to a Trusted Execution Environment.
- Item Tower: Deployed on the ad platform server.
- Input: Ad creative features, bidding information, etc.
- Output: Ad vector (Item Embedding), which can usually be pre-calculated offline and cached.
3. Matching and Ranking: Introducing TEE (Trusted Execution Environment)
After the User Embedding is generated, how can it be safely matched with massive Item Embeddings (calculating )? Here, TEE (such as Intel SGX, ARM TrustZone) is usually introduced to break the "Performance-Privacy" impossible triangle.
Architecture Flow:
1. The user side calculates (encrypted transmission).
2. The ad platform sends the candidate set (or index).
3. Both parties' data are decrypted and subjected to vector retrieval and scoring within the TEE (Confidential Computing Environment).
4. The TEE outputs only the Top-K ad IDs, without leaking specific vector values or user features.
This solution utilizes hardware-level memory isolation, with calculation speeds close to plaintext (performance loss is usually within 20%-30%), which is far superior to pure software cryptographic solutions.
4. Trade-off Strategy Between Training and Inference
When summarizing this architecture in an interview, one should clearly point out the fundamental differences in privacy strategies between Training and Inference:
Stage | Core Constraints | Common Tech Stack | Trade-off Points |
|---|---|---|---|
Offline Training | Model accuracy, data compliance | Vertical Federated Learning (VFL), MPC, Homomorphic Encryption | Heavy encryption, light real-time. To ensure secure gradient exchange, hour-level training duration is tolerable. |
Online Inference | Extreme low latency (<100ms) | Two-Tower Architecture, TEE (Confidential Computing), Local Differential Privacy | Heavy architecture, light encryption. Avoiding complex interaction protocols through physical isolation (client-side) or hardware trust (TEE). |
This hybrid mode of "Soft Core (Crypto) for Training, Hard Core (TEE/Architecture) for Inference" is currently the optimal engineering solution for balancing ad business monetization efficiency and user privacy compliance.
Trade-offs: The Game Between Accuracy and Privacy Protection

When discussing privacy computing in interviews, the easiest trap to fall into is claiming a "perfect solution." In reality, privacy computing is a zero-sum game process among Accuracy (Utility), Performance (Efficiency), and Privacy Security (Privacy). For advertising recommendation models, CTR (Click-Through Rate) or CVR (Conversion Rate) prediction models are extremely sensitive to accuracy; a tiny drop in AUC can lead to significant revenue loss. Therefore, candidates must be able to clearly articulate the "cost" under different technical routes.
Differential Privacy (DP): Direct Exchange of Noise and Accuracy
The core mechanism of Differential Privacy is injecting mathematically calibrated noise (such as Laplace noise or Gaussian noise) into data or gradients. In federated learning scenarios, this usually manifests as Local Differential Privacy (LDP) or Centralized Differential Privacy (CDP).
- Privacy Budget (): This is a key hyperparameter. The smaller the value, the stronger the privacy protection, the larger the injected noise, the slower the model convergence, and the lower the final accuracy; conversely, the larger the value, the smaller the noise, the higher the model utility, but the risk of privacy leakage increases.
- Pain points in advertising scenarios: Recommendation systems usually face sparse data problems, and Long-tail features are very sensitive to noise. If a smaller is set to meet strict compliance requirements, it may cause the model to fail to capture fine-grained user interests, directly dragging down online conversion effects.
- Trade-off strategy: In interviews, it should be pointed out that DP is more suitable for coarse-grained statistical analysis or as an auxiliary means, rather than the main defense means for fine-ranking models pursuing extreme precision. As relevant research points out, noise injection itself introduces inaccuracies, which may slow down convergence speed or reduce the accuracy of the final model.
MPC and HE: The Performance Cost of Lossless Training
Unlike differential privacy, Secure Multi-Party Computation (MPC) and Homomorphic Encryption (HE) can theoretically achieve Lossless training. That is, the model parameters trained based on these encryption protocols are consistent with the results trained on plaintext datasets.
However, this "lossless accuracy" comes at the cost of huge computation and communication overhead:
- Computation inflation: Homomorphic encryption involves large integer arithmetic and polynomial arithmetic. Compared with plaintext floating-point arithmetic, the calculation volume grows exponentially.
- Communication bottleneck: MPC protocols (such as Garbled Circuits or Secret Sharing) require frequent interaction between multiple parties during every round of gradient aggregation. According to test data from the 2022 China Privacy Computing Industry Research Report, the average time consumption of vertical federated learning modeling is tens or even hundreds of times slower than plaintext, and the gap will continue to widen as the data scale increases.
Interview Answering Strategy:
When asked "How to ensure model effect does not decline," do not make blind promises.
- Honest answer: "If we adopt strong security assumptions based on HE/MPC, we can achieve lossless model precision, but the training duration will extend from hours to days, which is unacceptable for advertising models that require hourly updates (Online Learning)."
- Compromise solution: The industry often adopts hybrid protocols or approximate computing. For example, only performing high-intensity encryption on core gradients, or using Partially Homomorphic Encryption (Paillier) instead of Fully Homomorphic Encryption, or even using plaintext transmission in non-sensitive layers, to exchange for usable training speed on the bottom line of legal compliance.
Summary: Decision Matrix for Technology Selection
When designing a system, you can refer to the following comparison dimensions to demonstrate your architectural thinking:
Dimension | Differential Privacy (DP) | Secure Multi-Party Computation (MPC) / Homomorphic Encryption | Federated Learning (No Encryption Enhancement) |
|---|---|---|---|
Model Accuracy | Lossy (Depends on budget) | Lossless (Theoretically equivalent to plaintext) | Lossless |
Training Performance | High (Only adds small noise generation overhead) | Low (Huge computation/communication overhead) | Medium (Limited by network bandwidth) |
Applicable Scenarios | Statistical reports, coarse ranking recall, user profile clustering | Offline model training, sample alignment (PSI) | Cross-institution joint modeling (Needs to be combined with encryption) |
Major Risks | Accuracy decline leads to revenue loss | Excessive time consumption leads to model update lag | Gradient Leakage attacks |
Excellent candidates can not only explain these technical principles but also combine them with business scenarios (e.g., during the "Double 11" promotion, the requirement for real-time performance is higher than privacy intensity fine-tuning) to provide strategies for dynamically adjusting privacy budgets or encryption intensity.







