In competitive interviews at top domestic private funds (e.g., High-Flyer, Ubiquant), the deciding factor is often not complex mathematical derivation or coding skills, but deep insight into A-share Market Microstructure. Many candidates with overseas backgrounds habitually transplant mature US T+0 high-frequency mean reversion or market-making logic, overlooking that A-share's unique T+1 system and price limits represent not merely parameter differences, but a fundamental reshaping of underlying factor logic. This lack of localization causes many strategies with perfect backtests to fail in live trading due to the inability to close positions intraday or liquidity gaps. True A-share quantitative factor mining requires transcending pure data fitting to deeply understand retail-dominated gaming characteristics, call auction information density, and the "magnet effect" of price limits. Interviewers focus on whether you can distinguish if Alpha stems from market inefficiencies or a misunderstanding of rules. Only by establishing a mining framework based on regulatory constraints can one avoid classic traps like "liquidity illusions" and demonstrate strategy robustness in the real Chinese market.
Core Differences: Why Copying US Stock Factors Leads to "Failing" A-Share Interviews?
In interviews with top-tier hedge funds (such as High-Flyer and JiuKun), the most common "failure point" is not the candidate's lack of mathematical derivation ability, but rather a lack of profound understanding of the A-share Market Microstructure. Many candidates with overseas backgrounds are accustomed to the continuous double auction mechanism and T+0 environment of US stocks, attempting to directly migrate mature "high-frequency reversal" or "market-making logic" to A-shares. However, this copying often leads to strategies performing excellently in backtesting but being completely unexecutable in live trading.
The unique trading system of A-shares is not just a difference in parameters; it fundamentally changes the underlying logic of factors. The core of what interviewers examine is: Are you clear whether your Alpha comes from market inefficiencies, or merely from a misunderstanding of the rules?
Core Comparison of Market Microstructure Between China and the US
Before delving into specific factors, a clear framework of institutional constraints must be established. The following are the most fatal differences between A-shares and US stocks when implementing quantitative strategies:
Dimension | US Equities | China A-Shares | Core Pitfalls in Quant Interviews |
|---|---|---|---|
Trading Mechanism | T+0 (Intraday turnaround) | T+1 (Buy today, sell tomorrow) | Pure intraday high-frequency reversal cannot form a closed loop; overnight risk must be borne. |
Price Limits | No hard individual stock limits (Circuit breakers only) | 10% / 20% / 30% Price limits (Limit Up/Down) | Liquidity dries up instantly after hitting the limit, leading to "liquidity illusions." |
Investor Structure | Institution-dominated (90%+) | High proportion of retail investors and hot money (Significant "80/20" phenomenon) | More Noise Trading; irrational volatility brings unique game-theory factors. |
Matching Mechanism | Continuous auction, fragmented liquidity across multiple exchanges | Call auction + Continuous auction, monopoly by single exchange | The call auctions at the close and open contain extremely high Alpha information density. |
Logic Failure Caused by Institutional Constraints
Given the differences mentioned above, two classic quantitative logics will fail directly in A-shares, which are also "minefields" that must be avoided in interviews:
- "Lock-in" Risk of High-Frequency Arbitrage
In a T+0 market, market makers or high-frequency strategies can complete "buy low, sell high" within millisecond time windows to earn the spread. But under the A-share T+1 system, a mispricing signal identified at 10:00 AM requires holding the position until the next day after buying. This means your prediction win rate must be able to cover overnight volatility risk. As pointed out by CSC Financial research, the trading structure of A-shares during the late trading session often exhibits characteristics of retail domination and short-termism; this concentrated exchange behavior of funds is largely to avoid the holding uncertainty brought by T+1. - "Liquidity Illusion" Caused by Price Limits
In US stocks, prices change until buying and selling forces balance. In A-shares, once the limit up is sealed, sell orders disappear, and volume drops sharply to zero. If your factor merely interprets "volume shrinkage" as "declining market attention" or "poor liquidity," you will reach a completely wrong conclusion. In fact, shrinking volume at the limit up represents extreme buying willingness (strong consensus expectation). This Liquidity Fracture caused by the system requires us to perform special cleaning or Dummy Variable processing on limit up/down states when processing data; otherwise, linear models will produce serious biases.
Understanding these hard constraints is the prerequisite for mining factors with "Chinese characteristics." Next, we will specifically explore how to dance in these shackles and mine Alpha that adapts to the A-share ecosystem.
The Impact of the T+1 Trading System on Factor Decay

When interviewing at top quantitative private equity firms like High-Flyer and Jiukun, a classic "killer" question is: "If you discover a minute-level reversal signal with an extremely high Sharpe Ratio, how would you deploy it in live trading?"
If you answer directly, "Monitor order book imbalance and place orders immediately for arbitrage," the interviewer might interrupt you right away—because you have ignored the most fundamental hard constraint of the A-share market: the T+1 trading system.
In US stock or cryptocurrency markets, the core of high-frequency trading (HFT) often lies in capturing millisecond or second-level price dislocations and quickly closing positions to lock in profits. However, in the A-share market, except for special scenarios involving inventory for T+0 enhancement, the vast majority of Alpha strategies must face a cruel reality: Stocks bought today (Day T) must be held overnight and cannot be sold until after the market opens tomorrow (Day T+1).
This mechanism fundamentally changes the logic of factor mining, mainly reflected in the "factor decay" risk across the following two dimensions:
1. Mismatch Risk of Signal Validity Period
The T+1 system forcibly lengthens the holding period. A strong buy signal issued at 10:00 AM, even if it accurately predicts the subsequent 30-minute rise, may be ineffective for A-share strategies.
- Scenario Deduction: Suppose your factor captures a rush of funds into a stock at 10:00, causing an instant price surge. After buying, you must bear all market volatility from 10:00 to the close (15:00), and then to the next day's opening (09:30).
- Result: If the stock experiences an intraday reversal and falls at 14:00, or if a crash in US stocks overnight causes A-shares to open lower the next day, your originally precise "30-minute prediction capability" will not only fail to be realized but will turn into a loss due to the forced overnight holding.
- Interview Response: When constructing factors, you must emphasize the timeliness of the prediction target. For A-share high-frequency factors, interviewers value whether you have tested the signal's predictive ability for the next day's opening price (Open_{t+1}) or the next day's average price (VWAP_{t+1}), rather than just the current period's return.
2. Shift in Mining Focus: Overnight Returns and Call Auction
Since T+1 locks up intraday liquidity, the focus of high-frequency factor mining in A-shares shifts from "intraday error correction" to "overnight gaming."
- Overnight Returns:
A significant amount of Alpha is actually generated during the non-trading period from the close to the next day's opening. Research by GF Securities points out that overnight returns (ret_overnight) and information during the opening call auction are highly characteristic factor sources in A-shares. In an interview, you can mention focusing on call auction data from 9:15-9:25, especially the order pressure during the irrevocable stage after 9:20, which often reflects the true intentions of major funds better than the intraday continuous auction.
> High-Frequency Data Factor Mining Based on Deep Learning mentions that utilizing the return rate of the opening price relative to the highest/lowest price during the call auction (such asret_open2AH1) can effectively capture early trading testing behavior by funds. - Tail-end Gaming and Intraday Structure:
Due to T+1 restrictions, retail investors and "hot money" tend to concentrate trading at the tail end (after 14:30) to reduce the time of uncertainty for overnight holding. Microstructure research by CSC finds that amplified volume at the tail end often implies intensified chip exchange; this heterogeneity in trading structure causes factor performance during the tail-end period to be distinctly different from that of early trading. - Practical Advice: Demonstrate your sensitivity to data segmentation in the interview. For example, you can propose slicing the full day of trading to separately construct "morning factors" (institution-dominated, good liquidity) and "tail factors" (retail-dominated, intense gaming), and point out decay characteristics based on indicators like Short-Term Trading Crowdedness (STC) in different time periods.
Summary: In an A-share interview, when you present a high-frequency factor, be sure to proactively add: "Considering T+1 constraints, I additionally tested the decay speed of this factor under overnight holding conditions and focused on analyzing its contribution to the next day's opening returns." This immediately demonstrates that you not only understand mathematics but also understand the microstructure of the Chinese market.
Liquidity Discontinuity Caused by Limit Up/Down

In A-share quant interviews, interviewers often test how candidates handle "extreme data." US stocks do not have price limits (except for circuit breaker mechanisms), so price discovery is continuous; whereas the A-share 10% (or 20%) limit rule artificially severs liquidity, causing prices to become ineffective the moment they hit the board. This phenomenon is called Liquidity Discontinuity.
If you merely answer "exclude limit up/down data" in an interview, you might be considered to lack a deep understanding of A-share microstructure. Top private funds value how you mine highly significant Alpha from these seemingly "invalid" data.
1. Magnet Effect and Order Book Gaming
Price limits are not just a static price boundary; they possess a significant "Magnet Effect." When the stock price approaches the limit up price, due to investors fearing they won't be able to buy (Fear of Missing Out), buy orders accelerate their inflow, causing the price to be acceleratedly "sucked" towards the limit board.
In factor mining, this is not just momentum, but a drastic change in Microstructure. You can mention the following logic in interviews:
- Order Piling and Cancellation Rate: On the eve of hitting the limit, the volume of Bid 1 orders and cancellation behavior in Level-2 data are key to predicting whether the board will be "sealed." If buy orders are frequently cancelled (spoofing) when approaching the limit price, it often indicates major players are luring longs; conversely, determined order piling is a strong signal of sealing the board.
- Signal Reversal of Liquidity Exhaustion: In normal trading, shrinking volume usually means declining attention. However, at the limit up, extremely shrinking volume (limit up on low volume) represents the strongest bullish sentiment—because there are no sell orders, buy orders cannot be executed. If you directly use standardized volume (Volume Z-score) when constructing price-volume factors, it will cause these strongest stocks to have extremely low scores. You must apply special tagging or non-linear processing to the "limit up state" in your model.
2. Mining Alpha from "Sealed Boards": More Than Just Ups and Downs
For factor mining regarding limit up/down, the core lies in quantifying the "hardness" of the board and its predictive power for next-day returns (especially overnight returns). The following are several specific mining dimensions, suitable for showcasing as technical cases in interviews:
- Bid Order Strength:
Calculate the ratio of the sealed order amount at the limit price to the total turnover of the day. The larger the ratio, the more the bullish intention exceeds the current liquidity supply, and the probability of a Gap Up the next day is extremely high.
> Formula Example: - Time Dimension Factors:
- Time to First Limit: Stocks that seal the board between 9:30 - 10:00 AM usually have a higher premium the next day than stocks that seal via a "sneak attack" at the close. Early sealing represents determined major capital, while late sealing is often hot money gambling on the next day's premium, which is prone to encountering the "nuclear button" (massive sell-off).
- Open Board Frequency: Count the number of times and duration the limit board opens intraday. Frequent opening usually implies huge divergence between longs and shorts, serving as a strong signal for reversal or high volatility.
- Overnight Return Prediction:
Under the A-share T+1 system, the core return of limit up strategies often comes from overnight gaps. GF Securities Research points out that pre-market price-volume information (such as overnight returnret_overnight) contains a lot of gaming information. You can build a prediction model specifically for limit-up stocks to predict the strength of the next day's call auction, thereby deciding whether to place an order to take profit or continue to hold the position.
3. Pitfall Avoidance Guide: Specifics of Data Processing
When answering questions about data cleaning, be sure to emphasize that you cannot simply exclude limit up data, otherwise serious Survivorship Bias will be introduced.
- Wrong Approach: Directly removing rows of limit up days when calculating Moving Averages (MA) or volatility.
- Correct Approach: Realize that the "true price" on a limit up day is actually higher than the displayed price (Shadow Price). When training machine learning models, you can input "is limit up/down" as an independent Categorical Feature, or use truncated regression methods like the Tobit model to correct the potential price.
By demonstrating control over these details, you can prove to the interviewer that you not only understand algorithms but also understand the unique trading rules and capital gaming logic of the A-share market.
Practical Insights: Three Major "Characteristic" Factor Mining Directions for A-Shares
When interviewing at top domestic private funds (such as High-Flyer, JiuKun, Lingjun), what interviewers value most is not your recitation of general factors (like momentum, value), but whether you understand the unique market microstructure and investor behavior of A-shares. Because the A-share market has a T+1 trading system, price limit restrictions, and an extremely high proportion of retail investors, directly copying the factor logic of US stocks often encounters severe "inadaptability."
To stand out in an interview, you need to demonstrate a deep understanding of the following three "Chinese characteristic" mining directions. This is not only the main source of excess returns (Alpha Source) for current quantitative institutions, but also the key dividing line between the "theoretical school" and the "practical school."
A-Share Characteristic Factor Mining Landscape
The following are the three core tracks that top domestic quantitative institutions are currently competing in. It is recommended to demonstrate your research framework around these directions during the interview:
Mining Direction | Core Data Source | Alpha Logic | Typical Factor Examples |
|---|---|---|---|
1. Market Microstructure | Level-2 High-Frequency Data<br>(Tick-by-tick transactions, Tick-by-tick orders) | Institutional Order Splitting and Front-running: Using millisecond-level data to identify traces left by institutional algorithmic trading (TWAP/VWAP), as well as fake orders during the morning call auction. | • Order Flow Imbalance<br>• Call Auction Cancellation Rate<br>• Minute-frequency Return Skewness (CSKEW) |
2. Behavioral Anomalies | Dragon and Tiger List (Longhubang)<br>Social Sentiment (Guba/Tieba) | Retail Herding Effect: A-share retail investors are easily driven by emotions to chase highs and cut lows. Using game theory data from specific seats (Hot Money vs. Institutions) to capture reversal signals after sentiment overheating. | • Hot Money Seat Premium Factor<br>• Retail Sentiment Reversal Factor<br>• Overnight Return |
3. Alternative Data | Analyst Reports<br>Interactive Platform Q&A | Information Asymmetry: Using NLP technology to analyze changes in analyst tone, or monitoring the response frequency of listed companies on interactive platforms to capture early leakage of fundamental information. | • Analyst Revision Sentiment Factor<br>• Interactive Platform Response Latency Factor |
In the following section, we will deeply deconstruct the two directions with the most practical value: Microstructure Mining based on Level-2 Data and Retail Sentiment Factors based on the Dragon and Tiger List, providing you with technical details directly applicable to interviews.
Order Flow Imbalance (OFI) Based on Level-2 Data

In interviews at top quantitative private equity firms, interviewers often skip basic price-volume factors (such as simple momentum or reversal) and directly assess the candidate's ability to process Level-2 high-frequency data. The unique microstructure of the A-share market makes factor mining based on Order Flow Imbalance (OFI) a key watershed distinguishing "academics" from "practitioners."
The core logic of OFI lies in capturing the "pre-execution" intent rather than the "post-execution" result. Traditional volume factors are lagging, whereas changes in the Order Book often contain leading Alpha information.
1. Mining "False Declarations" During the Call Auction Period
One of the most typical features of A-shares is the 9:15-9:25 Call Auction mechanism. The signal-to-noise ratio of data during this period is extremely high, making it a gold mine for mining high-frequency factors.
- Time Segmentation and Cancellation Gaming:
- 9:15-9:20 (Cancellable Phase): This is a peak period for major funds to engage in "luring longs" or "luring shorts." Smart Money often places huge buy orders to push up the virtual matching price, attracting retail investors to follow suit, and then instantly cancels the orders at 9:19:59.
- 9:20-9:25 (Non-Cancellable Phase): This is the real buying and selling game.
- Factor Construction Approach:
You need to construct a factor that measures "false pressure." For example, calculate the weighted buy order cancellation rate for the first phase (cancellable). If a stock sees a surge of buy orders before 9:19, but the buy volume drops off a cliff after entering 9:20 (i.e., major players cancel orders), this is usually a strong intraday Short Signal.
According to Quant Wiki's research notes, using fields such asret_open2AH1(return of the opening price relative to the highest price in the first phase) ordiverge_A1(amplitude in the first phase) can effectively quantify the intensity of this pre-market gaming.
2. Order Book Pressure During Continuous Auction
After entering the continuous auction, simple buy/sell snapshots are no longer sufficient to construct strong factors; you need to use the changes in Tick-level data to construct OFI.
- Basic OFI Formula:
Where represents the order quantity. A simple understanding is: Net increase in the Bid Book - Net increase in the Ask Book. - A-share Specific Weighted Price Fields:
In A-share L2 data,WeightedAvgBidPx(weighted average bid price) andWeightedAvgAskPx(weighted average ask price) contain deeper information than the best bid/ask prices. - Depth Imbalance Factor: When the stock price rises, but
WeightedAvgBidPxmoves downward (indicating that while buy orders are executed, pending orders are mainly concentrated at deeper levels, weakening support willingness), this divergence often predicts that the rise is unsustainable. - MPC-type Factors: Referring to CITIC Securities' research on high-frequency order imbalance, by calculating Market Participation Capability (MPC) and the skewness of order flow, one can more accurately predict micro-price trends.
- Depth Imbalance Factor: When the stock price rises, but
3. "Institutional vs. Retail" Order Tagging
This is a bonus point in interviews. The Shenzhen Stock Exchange's L2 data provides Order-by-Order data, while the Shanghai Stock Exchange mainly provides Snapshots. For Shenzhen stocks, you can reconstruct "large orders" and "small orders" using transaction-by-transaction data.
- Logical Inference:
- Retail Characteristics: Integer multiples of lots (e.g., 100 shares, 500 shares), smaller order amounts, and frequent placement at round number price levels.
- Institutional/Quant Characteristics: Non-integer lots (caused by algorithmic order splitting, e.g., 317 shares), extremely fast order placement speed, and placement positions often within the Spread between the best bid and best ask.
- Factor Construction:
Calculate the difference between Institutional Net Inflow (Inst_OFI) and Retail Net Inflow (Retail_OFI). Empirical experience shows that when RetailOFI is significantly positive (retail investors frantically placing buy orders) while InstOFI is negative, it is an excellent Reversal shorting opportunity.
Summary: When answering such interview questions, avoid only discussing generic "price-volume relationships." You must stick closely to Tick data field details (such as cancellation volume, weighted average price) and A-share specific time windows (the 9:20 cancellation deadline); this is the "microstructure cognition" that top private equity firms want to see.
Retail Sentiment Factors: Dragon and Tiger List & Forum Sentiment

In the A-share market, retail investors contribute an extremely significant amount of trading volume, which stands in sharp contrast to the institution-dominated US market. For top quantitative private equity firms like High-Flyer (Huanfang) and Juekun, quantifying "irrational behavior" is a crucial battlefield for mining Alpha. In an interview, if you can elaborate on "Chinese characteristic" behavioral finance factors from the two dimensions of Dragon and Tiger List (LHB) seat gaming and Guba (Stock Bar) sentiment NLP mining, you will be highly competitive.
1. Dragon and Tiger List Data Mining: "Smart Money" and "Leek Orders" Behind the Seats
The Dragon and Tiger List discloses the top five buying and selling seat data for stocks with abnormal movements (such as price deviation reaching 7%, excessive turnover rate, etc.) daily. Unlike the anonymous flow of US stocks, the Dragon and Tiger List directly exposes the attributes of the capital. When constructing factors, the core logic lies in Seat Labeling and Game Analysis of Capital Flow.
- Seat Labeling System:
- "Institutional Dedicated" & "Northbound Capital": These usually represent "Smart Money" driven by fundamentals. Research shows that net buying by institutional seats often possesses strong trend continuity.
- Well-known Hot Money (Youzi): For example, "Zhang Mengzhu", "Chao Gu Yang Jia" (Stock Trading Family Support), or specific brokerage branches (such as Caitong Hangzhou Shangtang Road). This type of capital is often aggressive in style, skilled at creating "consecutive limit-up" trends, but is also accompanied by high volatility and "pig butchering" (pump and dump) risks. When constructing factors, it is necessary to distinguish between "Strategic Hot Money" (locking positions to boost price) and "One-Day Tour Hot Money" (dumping the next day).
- Retail Base Camp (Lhasa Legion): Seats like East Money Lhasa Tuanjie Road are usually considered gathering places for retail investors. If the buying seats on the Dragon and Tiger List are dominated by the "Lhasa Legion," it usually implies that chips are loosening and the main force is exiting, which is a strong Reversal signal.
- Identification of Quant Private Equity Tracks:
This is an advanced interview topic. Since quantitative funds trade at high frequencies and are dispersed, it is traditionally difficult to capture them on the Dragon and Tiger List. However, by cross-validating top 10 circulating shareholders with Dragon and Tiger List branches, one can identify whether specific seats are "associated seats" of quantitative private equities. Once certain branches are identified as exhibiting long-term characteristics of quantitative capital (such as mechanical order placement, intraday rotation), their buying and selling behavior can serve as a special "peer capital" factor.
2. Alternative Data: NLP Sentiment Factors from Guba and Forums
A-share retail investors rely heavily on community interaction. East Money Guba and Xueqiu are the core battlegrounds for sentiment fermentation. Compared to Twitter/Reddit for US stocks, the correlation between discussions in domestic stock bars and stock price movements is more significant in small-cap stocks.
- Retail Overheat as Reversal:
The classic logic is a Contrarian Indicator. When the discussion heat (Buzz) for a certain stock in Guba suddenly spikes, and the sentiment is extremely high (screen full of "Good News", "Limit Up"), it is often a signal of a local peak. - Factor Construction Example: Calculate
(Current Post Count - Past N Days Mean) / Past N Days Std Dev. When this Z-score is greater than a specific threshold (e.g., 2.0) and the stock price is at a high level, the short selling signal is significant.
- Factor Construction Example: Calculate
- NLP Text Mining Details:
Mentioning specific technical details in an interview adds points. For example, simple Bag of Words statistics have limited effectiveness in the Chinese context and need to be combined with pre-trained models like BERT for sentiment classification. - Key Features: Besides the "Long/Short" ratio, Disagreement is also an important factor. When there is intense mutual abuse between bulls and bears in posts and emotional divergence is extreme, it is often accompanied by an amplification of trading volume and an increase in volatility, making it suitable for constructing volatility strategy factors.
3. Pitfall Guide: Data Cleaning and Adversarial Tactics
When using the above data, you must demonstrate to the interviewer your awareness of data noise:
- "Water Army" (Spam Bot) Identification: There are a large number of bots or "pig butchering" guide posts in Guba. Noise needs to be removed through methods like posting time distribution (e.g., concentrated posting late at night) and IP address clustering.
- Seat Vest Changing: Hot money brokerage branches change frequently. The factor decay of a single seat ID is fast, requiring dynamic maintenance of a "Seat Pool".
Summary: Mining sentiment factors in A-shares is essentially using data to find the critical point of "irrational exuberance". Whether it is the seat gaming on the Dragon and Tiger List or the overheating of public opinion in Guba, the core is to utilize the retail herding effect for contrarian operations or liquidity provision.
Deconstructing the Interview Question Bank: How to Showcase Your "Mining Framework"?
When facing interviews at top private funds like High-Flyer (Huanfang) and Jiukun, you will often encounter a type of open-ended question: "Please design a factor to measure market sentiment" or "How would you mine a short-term price-volume factor?"
Junior candidates often rush to throw out specific formulas (such as Return / Volatility), whereas senior interviewers value your underlying industrialized mining process. From the perspective of top quantitative institutions, the Alpha of a single factor is transient, but a robust, iterative mining framework is the core competitiveness.
In an interview, it is recommended to use the following five-step standard process to answer such questions, demonstrating your complete closed-loop ability from "logical hypothesis" to "live trading implementation":
- Logic Hypothesis
Do not start with blind "brute force mining." First, articulate your economic intuition. For example, when constructing a reversal factor, is it based on "overreaction" or "liquidity compensation"? Even when using data mining techniques, demonstrating your understanding of market microstructure (such as the behavior of retail investors chasing rallies and selling in panic) proves that you are not simply fitting data. - Data Cleaning & Pre-processing
This is the key to distinguishing "Kaggle players" from "Real-world Quants." You need to actively mention how to handle data noise specific to A-shares, such as handling suspensions and resumptions, removing ST stocks, and adjustments for ex-rights and ex-dividends. Interviewers are very concerned about whether you are aware of the risk of Look-ahead Bias, such as erroneously using data that is only available after the close when calculating factors. - Formula Construction
This step transforms logic into mathematical expression. You can mention two paths:
- Logic-driven: Manually construct explicit formulas (like WorldQuant Alpha101 style), emphasizing the interpretability of the factor.
- Algorithm-driven: Use Genetic Programming or Machine Learning models to automatically generate non-linear factors. When mentioning this, be sure to balance the risk of a "black box" against mining efficiency.
- Backtest & Validation
Do not just talk about annualized return. A professional answer should cover multi-dimensional evaluation metrics:
- IC (Information Coefficient) & Rank IC: Core metrics for measuring predictive ability.
- ICIR: Evaluates the stability of the factor.
- Turnover: High turnover means high costs; actual performance after fees must be considered.
- Decay Test: Does the factor's performance decline rapidly out-of-sample?
- Risk Neutralization
Finally, demonstrate your risk control awareness. A raw factor often contains significant industry exposure or market cap exposure. You need to explain how to use Orthogonalization to remove the influence of industries (e.g., Shenwan Level 1 Industries) and Market Cap, ensuring that the mined Alpha is pure excess return, rather than Beta assuming some style risk.
The "Red Line" from the Interviewer's Perspective:
When listening to this framework, the interviewer is not only looking for the completeness of the steps but is also looking for your ability to handle Edge Cases. For example, when a factor fails in a specific year (such as the large-cap rally in 2017), your framework must not only be able to detect it but also explain the reason. Remember, demonstrating a "flawed but logically rigorous" process is far more likely to win an Offer than presenting a "perfect but inexplicable" Sharpe ratio.
Next, we will delve into the two most technically challenging parts of this framework: how to use genetic programming for automated mining, and how to handle those headache-inducing data special cases in A-shares.
Applications and Pitfalls of Genetic Programming in Mining

In interviews with top private funds (such as High-Flyer and JiuKun), Genetic Programming (GP) is not just a technical term, but a core testing point for a candidate's ability to handle "nonlinear mining" and "automated factor production." Interviewers usually won't ask you to write genetic algorithm code by hand, but will examine your deep understanding of Operator Trees, Fitness Functions, and Overfitting Control.
1. Core Concepts: From Manual Logic to Operator Evolution
You need to not only explain how GP automatically generates factors by simulating biological evolution (selection, crossover, mutation) but also emphasize its specific form in quantitative finance.
- Operator Tree Structure: Explain how you construct factor expressions. For example, leaf nodes are basic data (Open, Close, Volume), and internal nodes are function operators (
ts_rank,correlation,delay). - Fitness Function: A common interview question is "What is your optimization objective?". Besides the common IC (Information Coefficient) or IR (Information Ratio), high-level answers can mention penalties for turnover rates, or adding constraints on formula complexity within the fitness function.
2. Interview Must-Ask: How to Prevent Overfitting?
The biggest trap of GP is that it easily generates a complex formula that "perfectly fits historical data" but fails out-of-sample. When an interviewer asks "how to prevent overfitting," avoid vague talk and provide specific engineering solutions:
- Complexity Penalty: Explicitly state adding a penalty term for formula length (tree depth or node count) in the fitness function. As shown in AlphaForge's research, factor expression lengths tend to reach the upper limit, and excessive pursuit of complex mathematical combinations often leads to out-of-sample decay.
- Strict OOS (Out-of-Sample) Testing: Emphasize the time-series isolation of "Training Set - Validation Set - Test Set" and mention using Rolling Windows to detect the temporal stability of factors.
- Adversarial Validation: Mention adding random noise to training data to observe if factor performance fluctuates drastically, thereby eliminating factors that "cheat" by exploiting tiny data noises.
3. Avoiding Pitfalls: "White Box" Means More Than Just Visible Formulas
Although factors generated by GP have explicit formula forms compared to neural networks and theoretically belong to "white box" models, directly showing a 20-line nested formula in an interview will not score points. Interviewers are extremely wary of "black box" thinking where one "mines for the sake of mining."
High-Score Answer Strategy:
- Economic Rationale: Demonstrate your ability to perform "logical attribution" on machine-generated factors. For example, if GP generates
rank(close - delay(close, 5)) / volume, don't just say it backtests well; explain it as "a liquidity-adjusted short-term momentum factor." - Operator Pruning: Describe how you simplify formulas. For instance, if you find a subtree contributes negligibly to IC but increases complexity, you would manually or automatically prune that branch.
- Avoid Over-mining: Emphasize that you limit the operator set. For example, when dealing with Alpha101 type price-volume factors, be cautious about using high-order statistical moments (like skewness, kurtosis) because they are extremely sensitive to outliers and can easily cause the model to lose control under extreme market conditions.
Summary: In an interview, answers regarding Genetic Programming should convey a sense of balance—you possess efficient means for automated mining, but also maintain the risk control awareness of subjective quantitative analysis, and do not blindly trust complex mathematical coincidences generated by machines.
"A-Share Exceptions" in Data Cleaning: Suspension, Ex-Rights/Ex-Dividend, and ST
When interviewing with top private funds (such as High-Flyer, JiuKun), interviewers often ask about the details of data cleaning to judge whether a candidate has only run cleaned datasets on Kaggle or has truly handled the complex "dirty data" of A-shares. The unique Market Microstructure of A-shares causes standard US stock cleaning logic to often fail here. The following are three "Chinese characteristics" data pitfalls that must be mastered.
1. Suspension: Liquidity Black Holes and "Catch-up Gains/Losses"
Historically, A-shares have experienced large-scale, long-cycle arbitrary suspensions (although this has improved in recent years, it still exists). In a backtesting engine, simply "deleting missing values" is often insufficient.
- Trap Scenario: A stock is suspended for 3 months, during which the broader market rises by 20%. On the day of resumption, the stock hits "limit-up at the open" (一字涨停) for 5 consecutive days to catch up.
- Backtest Distortion: If your strategy issued a "buy" signal during the suspension period, and the backtesting system defaults to executing at the "previous close" or "open price," your backtest curve will show a huge fake profit. In reality, capital could not have bought in at all.
- Processing Framework:
- Universe Construction: When building the daily tradable stock pool (Universe), stocks suspended on that day must be strictly excluded.
- Position Handling: For held stocks that suddenly suspend, the backtesting logic should force a Lock Position until resumption. During this period, you cannot simulate based on index returns; you must bear the risk of liquidity loss.
2. Ex-Rights & Ex-Dividend (Splits & Dividends): The "Look-Ahead" Risk of Adjustment Factors
The frequency of high stock dividends (stock splits) and cash dividends in A-shares is much higher than in US stocks. Handling price gaps usually involves Backward Adjustment, but the devil is in the details during factor mining.
- Adjustment Trap: Directly using backward-adjusted prices to calculate factors (e.g.,
Closeadj / Closeadjdelay1 - 1) is usually fine, but errors easily occur in factors involving absolute price values. For example, some logic relies on "share price below 5 yuan" as a filter for junk stocks. If backward-adjusted prices are used, Kweichow Moutai from ten years ago might appear as only a few yuan, leading to it being incorrectly filtered out. - Practical Solutions:
- Signal Generation: When calculating technical indicators (like MA, RSI) or returns, you must use backward-adjusted prices to maintain the continuity of the price series and eliminate the impact of gaps.
- Execution & Filtering: When dealing with transaction amounts, order prices, market capitalization filtering (Market Cap), or price threshold judgments, you must use the unadjusted Raw Price.
- Interview Bonus Point: Mention Point-in-Time (PIT) data. Ordinary adjustment factors are based on the "current" perspective. If a listed company revises its dividend plan, historical adjustment factors might change. Rigorous backtesting should use the adjustment factors "known at that time."
3. ST and Limit Up/Down: Survivorship Bias Under Hard Constraints
The ST (Special Treatment) system and the price limit (limit up/down) system in A-shares are hard constraints completely different from US stocks.
- ST Stock Processing: ST stocks not only face delisting risks, but their price limits are usually narrowed to 5%.
- Strategy: Most quantitative private funds directly exclude ST and \*ST stocks during the preprocessing stage when mining Alpha factors. In an interview, you should clearly state: retaining ST stocks introduces extreme Idiosyncratic Risk, and their liquidity often cannot support institutional capital.
- Limit "Hard Constraints":
- Buy Restrictions: If a factor predicts a stock will surge tomorrow, but the stock opens "limit-up" tomorrow, you actually cannot buy in.
- Backtest Correction: When calculating IC (Information Coefficient) or backtest returns, samples that cannot be bought (limit up) or sold (limit down) on the day must be excluded.
- Data Source Reference: In practice, you can use the
pricelimitstatusfield in thecnstockstatustable (as mentioned in the BigQuant Data Documentation, where status 1 is limit down, 3 is limit up) for precise filtering, rather than simply comparingCloseandHigh.
Summary Script:
When summarizing this section in an interview, you can phrase it like this: "In my mining framework, data cleaning is not just denoising, but a digital mapping of trading rules. I prioritize excluding ST stocks, suspended stocks, and new stocks listed for less than 60 days; I use adjusted prices when calculating factors, but strictly compare raw prices with limit up/down statuses when judging execution feasibility, ensuring that backtest returns are real returns that are 'attainable'."
A Guide to Avoiding Pitfalls: The "Pseudo-Logic" Interviewers Dislike Most
In interviews with top private funds like High-Flyer (Huanfang) and JiuKun, interviewers are often not worried about your mathematical derivation ability or code implementation ability—these are basic thresholds. Their keenest sense is used to capture your lack of "common sense" regarding the A-share market. Many candidates with shiny overseas backgrounds often step into the minefield of "pseudo-logic" because they directly apply US stock experiences or rely overly on pure mathematical mining.
The following are three major "red lines" that must be strictly avoided in interviews, which often signal that a candidate lacks practical experience or lacks reverence for the Chinese market.
1. Ignoring Transaction Friction: The Disappearing "High Sharpe"
The most typical rookie mistake is talking eloquently with a high-frequency strategy backtest report that hasn't deducted fees. In the US market, due to the market maker system and exchange Rebate mechanisms, transaction costs for certain high-frequency strategies are extremely low, and one can even gain revenue by providing liquidity.
But in the A-share market, transaction costs are the red line of life and death for a strategy.
- Stamp Duty: A-shares levy stamp duty on the selling side (historically mostly 0.1%, although there are policy adjustments, it remains a hard cost), which is a huge loss for intraday high-turnover strategies.
- Impact Cost: Top private funds have huge capital volumes. Your strategy might perform perfectly with a few million in capital, but with several hundred million, the buying and selling behavior itself will significantly push up or suppress the stock price, leading to a sharp increase in slippage.
Avoidance Advice: When demonstrating any high-turnover (e.g., daily turnover > 20%) factors or strategies, you must actively mention your assumptions regarding transaction costs. If your logic is "earning meager spreads through high-frequency mean reversion" but you haven't considered the high-fee environment of A-shares, the interviewer will directly judge the strategy as "unimplementable".
2. Parameter Over-optimization: Mining for the Sake of Mining
As artificial intelligence becomes mainstream in quantitative investment, many candidates like to demonstrate complex factors mined based on Genetic Programming or deep learning. These factor formulas may be three lines long, containing a dozen parameters, with backtest curves as smooth as a straight line.
In the eyes of interviewers, this is often equivalent to Overfitting.
- Data Mining Bias: If you tried 10,000 parameter combinations and finally selected the best performing one, that's not Alpha, that's luck.
- Lack of Logical Support: The interviewer will ask: "Why is this factor effective?" If you can only answer "because the data looks good" without explaining the underlying economic principles (such as liquidity premium, retail disposition effect, etc.), this will be regarded as pseudo-logic.
Avoidance Advice: Adhere to "Occam's Razor". The core logic of an effective A-share factor is often concise. Rather than showing a "black box" model with complex parameters, it is better to deeply analyze the capital gaming intent behind a simple logic (such as "rapid pull-up at the close").
3. Blindly Applying US Stock Logic: The "T+0" Mindset That Is Ill-Suited
Many candidates returning from Wall Street or studying US stock textbooks habitually apply US mid-frequency logic directly to A-shares, ignoring the essential differences in the microstructure of the two.
- T+1 vs. T+0: US stocks allow unlimited intraday rotation, suitable for intraday ultra-high-frequency mean reversion. A-shares implement a T+1 system; stocks bought cannot be sold on the same day (unless doing T0 through existing positions). If the strategy you designed relies on "selling for profit within 10 minutes of buying", in A-shares this usually requires securities lending or existing position support, and the cost and difficulty of securities lending are extremely high.
- Price Limit Constraints: US stocks have no price limits, while A-shares have a 10% (or 20%) price limit. Price limits lock up liquidity, causing buy orders to fail to execute or sell orders to fail to escape.
Avoidance Advice: When explaining strategies, you must make "localization" corrections. For example, when discussing reversal factors, explain how to handle the overnight holding risk brought by T+1, or how to utilize microstructure features like order cancellation behavior patterns to optimize execution, rather than assuming the market possesses infinite liquidity.
Summary: Economic Intuition > Complex Mathematics
The ultimate secret of the interview lies in: Do not sacrifice economic intuition for the sake of showing off mathematical skills.
Top private funds are not looking for "alchemists" who only know how to call packages and run data, but observers who can see the essence of market gaming through data. When you are mining "Chinese characteristic" factors, please always ask yourself: Who is the source of this factor's return? Is it retail investors chasing highs and cutting losses, or institutions forced to rebalance? Your mathematical model only has value if the logic holds water.







