Generated by Quarry-LDR on 2026-08-04: run 1497b2907a55, three search iterations, a 57,815-token cached evidence corpus, 10 sections, $2.88 of API spend. Cost anatomy in FirstRunReport.md. Verbatim model output except em dashes normalized to commas per repo convention and one duplicated section heading removed.
In OTC corporate bond markets, where a dealer's quoting policy itself reshapes the flow, spreads, and liquidity it will later be trained on, can a policy and the market response it induces be learned jointly as a self-consistent fixed point, and does that fixed-point policy generalize better than one trained under the standard assumption that the data-generating process is exogenous?
Corporate bond trading remains predominantly over-the-counter and dealer-intermediated, and its electronification has proceeded mainly through multi-dealer-to-client (MD2C) request-for-quote (RFQ) protocols rather than centralized limit-order books [1]. RFQ is described as the most widely used electronic execution protocol in bond markets [2], and MD2C platforms have become "the dominant architecture for institutional bond trading," letting clients solicit several dealers simultaneously and trade at the best price within a short response window [4]. Feedback is only partial: the winner learns the cover price, while losers typically learn only their ranking and whether a trade occurred, sometimes not even that [4]. For the dealer, quoting is therefore a joint problem of inventory management, counterparty selection, and endogenous execution probability: quote competitively enough to win flow, but not so aggressively as to destroy margin or accumulate unwanted inventory [1][3].
The central finding across the corpus is that execution probability is not an exogenous statistic. Hit ratio is "an equilibrium object shaped by platform design, client search, and rival participation" [1][6], and hit probabilities depend on the competitive set and on the number of solicited dealers, making "how many dealers are asked" a first-order primitive of RFQ economics [5]. Aggregator/routing layers add a second channel: even absent last look, "the aggregator itself shapes adverse selection and incentives" [5], and platform routing driven by slowly varying dealer win scores can generate fold bifurcations, bistability and hysteresis in score dynamics, producing an endogenous "campaign vs. harvest" pattern in optimal quoting [8].
This structure matches the formal framework of performative prediction, in which deploying a model shifts the data distribution it is evaluated on [15]. Its two solution concepts are performative stability, a fixed point where the model is optimal for the distribution it induces, so retraining returns the same parameters [11][14], and the performative optimum, which globally minimizes performative risk [12][14]. These generally differ: stable points need not minimize the moving-target risk, and the stability–optimality gap is generally nonzero [12]. Retraining (repeated risk minimization) is reinterpreted not as a nuisance but as the natural equilibrating dynamic whose fixed points are stable points [16][17]; it converges linearly under strong convexity, smoothness, and sufficiently Lipschitz (low-sensitivity) distribution maps [12][13][18][19], and becomes computationally hard (PPAD-complete) when performative effects are strong [12].
-
Control/equilibrium models. OTC market-making models built on stochastic control and Hamilton–Jacobi–Bellman equations handle size ladders and inventory risk [5][6], with extensions for hedging and market impact [21][31] and explicit hit-ratio targeting objectives [34]. However, most of this literature treats execution probability in reduced form, quote-dependent intensities or fill probabilities, "rather than elevating hit ratio itself to an explicit control target" [21]. Strategic settings are handled via Stackelberg or Nash equilibrium characterisations of competitive liquidity provision, and via reinforcement-learning/agent-based studies of tacit coordination on MD2C-type platforms [22][23][24]. Notably, Stackelberg equilibria in strategic classification are shown to coincide exactly with performative optima [25], supplying the formal bridge between game-theoretic dealer models and performative learning.
-
Causal identification from dealer data. Because dealers set spreads using market, RFQ, bond, client and competitive information, historical spread–outcome correlations "may not reflect causal relationships" [7][30]; estimating optimal spreads requires the interventional distribution (do-operator), not the historical conditional [10]. Graphical models plus the back-door criterion identify a minimal sufficient conditioning set, volatility plus RFQ, bond and client features [7], and this framework has been applied to optimal pricing, revenue-potential estimation, and axe–client matching, with the finding that neglecting confounders or partial observability (e.g., information asymmetry, price-discovery intent) biases models and yields suboptimal decisions [3][27][29][30]. Where identification fails, randomized controlled trials / A/B tests are the stated remedy [40]. Causal interventions have also been used for counterfactual pricing and revenue analysis on MD2C platforms, bridging structural RFQ models with discriminative learning [5][11-note: see 26], with Double Machine Learning applied more broadly in finance [26].
-
Estimation and evaluation. Discriminative models often outperform generative ones on predictive tasks and benefit from regularization and ensembling for out-of-sample generalization [38], with feature choice guided by the back-door minimal conditioning set plus domain knowledge [33]. Evaluation is hard precisely because of performativity: changed fill-probability estimates alter trading behaviour and thus provoke "unknown responses from the market that influence the conditions for the next time instance," so gains can only be assessed indirectly via backtesting with a fixed test environment and strict temporal train/test decoupling [39][41]. Recent learning-theory work makes this tension explicit: under performativity the learning target drifts with the prediction, and "the more a model is used to intervene..., the more the sample deviates from the original population" [35]; generalization bounds are obtained via covering numbers and Wasserstein sensitivity [36], with worst cases characterized as min-max (population self-negation, DRO) and min-min (sample self-fulfilment, DFO) functionals [37].
A related degenerative risk is model collapse: models trained on outputs of prior versions of themselves or on uncurated synthetic data progressively "forget the true underlying data distribution," first losing distribution tails and then converging to low-variance point estimates [42][44][45], driven by function-approximation, sampling and learning errors that compound in complex models [43][46]. Practically, dealer systems also operate under model-risk and regulatory constraints, MiFID II best-execution evidence, pre-trade price comparison and post-trade TCA [47][48][49], expanded ATS/TRACE reporting for fixed income [51], and at least one approach argues for enhancing input features rather than the learning algorithm precisely to keep model complexity manageable under such controls [50].
The corpus documents endogeneity in RFQ markets [1][5][8] and the performative-prediction toolkit [11–19][35–37] separately, but contains no study that explicitly applies performative stability/optimality or fixed-point retraining to bond RFQ auto-quoting, and no empirical estimate of the stability–optimality gap or of sensitivity/Lipschitz constants in this market. Evidence on whether performative-aware policies outperform standard retraining in live dealer settings is likewise absent; claim [39] explicitly notes such gains can only be assessed indirectly [39]. Similarly, [31]'s stated transient-vs-permanent impact distinction is asserted at claim level without supporting detail in the cited reference list, and should be treated as weak.
RFQ is the dominant mechanism, and it is a dealer-selection mechanism as much as a pricing mechanism. Corporate bond trading remains predominantly over-the-counter and dealer-intermediated even after substantial electronification; electronification has proceeded less through fully centralized limit-order books than through multi-dealer-to-client (MD2C) request-for-quote protocols, in which a buy-side client solicits a small panel of dealers and typically trades, if at all, with the best respondent [1]. Vendor-side descriptions agree that RFQ is the most widely used electronic execution protocol in bond markets: a trader sends a request to one or more dealers, who respond with a price within a specified response window [2]. MD2C platforms are described as the dominant architecture for institutional bond trading, allowing clients to simultaneously solicit quotes from multiple dealers and, within a short response window, choose whether to trade at the best available price [4]. The economic rationale offered for this structure is instrument proliferation: a single issuer may maintain hundreds of distinct securities with varying maturities, coupons and seniorities, which fragments liquidity and reduces the probability of matching counterparties through centralized order books [4].
Other protocols coexist with RFQ. A modern fixed-income EMS supports RFQ, click-to-trade, all-to-all (A2A), portfolio trading and algorithmic execution, and helps traders select the protocol per trade [47]; portfolio trading is a distinct protocol in which a firm sends an entire basket of bonds, potentially hundreds of instruments, to a dealer for simultaneous execution, useful for large-scale rebalancing, transitions and index-tracking [2]. The corpus documents that these protocols exist and are supported operationally [2][47], but it does not contain a mechanical description of continuous streaming quote provision in credit (e.g., stream lifetimes, tiering of streamed ladders, or last-look in bonds); the only last-look material concerns FX, where a short post-trade window in which a liquidity provider may accept or reject affects both spreads and selection [5]. Claims about streaming mechanics in corporate bonds should therefore be treated as unsupported by this corpus. Regulators observe the same protocol drift: FINRA reported increased trading of TRACE-eligible securities in dark pools, greater use of request-for-quote processes, and more executions of customer orders by firms acting as both broker and dealer [51]. On scale, roughly 50% of US investment-grade corporate bond volume is now executed electronically, up from about one-third in 2020 [49].
Sizing and multi-tier structure matter within a single RFQ. In OTC RFQ markets an additional structural feature is size heterogeneity: a liquidity provider typically answers a menu of quotes for different notionals, which leads naturally to multi-size ladder controls [5]. Correspondingly, formal models have the dealer quoting on both sides of each bond–client-tier pair across a size ladder with deterministic sizes, with RFQ opportunities arriving at an exogenous per-(bond, tier, side, size) intensity and the fill probability being a function of the quoted offset from mid, so that the realized OTC fill intensity is arrival intensity multiplied by fill probability [34].
The RFQ process is only partially observable, and the feedback structure is itself a design variable. In the MD2C description of [4], the winning dealer is informed of the second-best quote (the "cover price"), while other dealers receive feedback on their ranking and on whether a trade occurred, with the explicit caveat that in some cases the losing dealers cannot even distinguish a missed RFQ (trade occurred elsewhere) from a passed RFQ (no trade) [4]. A competing characterization from the market-making/ML literature is harsher: the practical setting is described as a blind auction with a hidden distribution of trade inquiries to selected participants, response windows on the order of seconds, lack of response-price transparency among competing bidders, and limited information about acceptance or rejection of submitted offers [32], with the process framed as algorithmic responses to RFQs conditional on the dealer's inventory interest and trading strategy [28]. These two accounts conflict on the granularity of post-RFQ feedback (cover price and ranking versus effectively blind), and the corpus does not resolve which is typical; both are cited here. Either way, the censoring is asymmetric: outcomes on unwon RFQs are observed coarsely or not at all [4][32], which is precisely the regime in which fill-probability estimation is described as "a key problem in this business" [28].
Decision variables documented in the corpus:
- Price/spread offsets from mid, per bond, per client tier, per side, per size rung of the ladder [34][5]; equivalently the "normalized half-spread," which adjusts the dealer's quoted spread by a market liquidity proxy such as half the bid-ask spread of a composite benchmark [33].
- Whether to respond at all / how aggressively, trading off hit probability against margin, adverse selection and post-trade market movements [3][4].
- Inventory and risk-transfer choices, including internalization versus external risk transfer, since dealer-market models have been extended to include hedging and market impact to capture that transition [21]; formal objectives include a running inventory-risk coefficient and a terminal inventory penalty [34].
- Client/counterparty targeting and commercial actions, e.g., identifying clients receptive to trading the dealer's axes, pre-existing positions the dealer wishes to buy or sell, with pricing skewed in favor of trading when the dealer has an axe [3][30].
- Explicit KPI targets: models now let the dealer set target hit-ratio levels per targeted client tier, with hit-ratio penalty weights alongside inventory-risk terms [34].
Observable outcomes:
- Hit ratio / win rate, including a size-weighted instantaneous hit ratio defined against a notional-arrival scale [34]; in industry practice, historical hit ratios are used to select dealers and route RFQs, especially for liquid electronic bond orders, alongside axes, response quality and relationship information [21]. Fermanian et al. are credited with estimating dealer-quote and client-reservation distributions and explicitly deriving ex-ante hit-ratio functions [1].
- Ranking and cover price where disclosed [4].
- Revenue measured at multiple horizons: end-of-day flow value using market mid-prices to value inventories (optionally with liquidity penalties for unrealized mark-to-market profit), and short-term flow value over a fixed post-trade window (seconds to minutes), the latter explicitly useful for identifying clients trading with information asymmetry before the signal is overridden by external market factors [29].
- Post-trade TCA / best-execution evidence: the EMS enables measurement against benchmarks, understanding of dealer relationships, and pre-trade side-by-side price comparison across counterparties [47], with MiFID II requiring firms to take sufficient steps to achieve best execution and retain evidence [48][49].
This is where the corpus is structurally rich but empirically thin. It supports the endogeneity claim through several distinct channels, but contains no reported effect magnitudes from an identified experiment or quasi-experiment in corporate bonds. That gap should be stated plainly.
Channel 1, the price→fill link is interventional, not associative. The optimal-pricing problem is written with a do-operator on the spread precisely because the dealer needs the interventional distribution of winning, not the historical conditional distribution of RFQs won given spreads; this distinction "requires a causal, rather than purely associative, relationship between the variables" [10]. The reason is behavioral: dealers do not quote based solely on external, context-independent drivers, but actively incorporate the market environment, the RFQ itself, instrument characteristics, client identity and the competitive landscape, which "introduces correlations between historical spreads and RfQ outcomes that may not reflect causal relationships" [7]. The graphical model makes the confounders explicit: because client identity is known on MD2C platforms, pricing policies incorporate client features, bond features such as liquidity conditions or relative value, and RFQ features such as the number of dealers in competition; ignoring these "can bias optimal pricing models, as these variables introduce spurious dependencies between the quoted price and the RfQ outcome" [30]. Remedies proposed are back-door adjustment on a minimal conditioning set (volatility, RFQ features, bond features, client features) [7], and, where effects are not identifiable from history, randomized controlled trials implemented as A/B tests [40].
Channel 2, the competitive set and win probability are equilibrium objects, not exogenous statistics. Synthesizing the RFQ literature, [1] concludes that "the dealer's hit ratio is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation," and that execution is shaped not only by price but by the protocol, the number of requested dealers, competitor response behavior, and the client's search technology [1]. Structural and empirical work on MD2C platforms models hit probabilities as depending on the competitive set and the number of solicited dealers, making "how many dealers are asked" a first-order primitive of RFQ economics [5]. Notably, [21] observes that most of the optimal-market-making literature still treats execution probability in reduced form (quote-dependent intensities or fill probabilities) rather than elevating hit ratio to a control target, i.e., the endogeneity is acknowledged conceptually more often than it is modeled as a control object [21].
Channel 3, a dealer's own aggressiveness feeds back into future access to flow. This is the sharpest endogeneity claim in the corpus. In aggregator-routed RFQ markets, two facts are asserted as empirical regularities: (i) the competitive set is typically a subset of eligible LPs, shaped by ranking, relationship, or platform "top list + exploration" logic; and (ii) inclusion and rank depend on slowly varying performance scores (e.g., long-run win ratio, response quality) "which in turn depend on how aggressively the LP quotes" [24]. Formalizing this, [8] models aggregator flow whose opportunity intensity is multiplied by a promotion gate driven by the dealer's win score, with wins updating the score, and shows that for steep (logistic) promotion gates the score dynamics can exhibit fold bifurcations, bistability and hysteresis, producing an endogenous "campaign versus harvest" pattern in optimal quoting [8]; background (non-gated) flow plays a stabilizing role in maintaining inventory-mixing capacity when the dealer is weakly promoted [8]. The corpus also documents the bond-market analogue of the routing rule itself: historical hit ratios are used to select dealers and route RFQs [21]. Caveat: [8] and [24] are FX/aggregator-focused; the corpus does not provide direct evidence that bond-platform shortlisting exhibits the same bistability, and [8]'s bifurcation results are stated as model and numerical-experiment outcomes rather than as estimated market facts [8].
Channel 4, adverse selection is a function of competition and of winning itself. Quoting too aggressively may increase hit probability but reduce margins or expose the dealer to adverse selection and post-trade market movements [3][4]. Oomen's aggregator argument is that the "winner's curse" can strengthen as more LPs compete, because the best displayed quote is more likely to be selected precisely in states unfavorable to the winning LP, which supplies an economic rationale for limiting the number of LPs exposed to each RFQ and for using routing and ranking rules rather than always inviting the full pool [23][1][6]. In other words, realized (as opposed to quoted) spreads and realized P&L are selected samples conditional on winning, and the selection intensity depends on the competitive configuration [23]. The corpus's suggested measurement device is the short-horizon flow value, which is designed to surface information-asymmetric clients before external factors dominate [29].
Channel 5, evaluation itself is contaminated. A backtesting study of bond RFQ fill-probability models states the endogeneity problem operationally: any performance gain from embedding a better fill-probability estimator "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs" [39]. It also notes that strict temporal decoupling of train and test sets means no data-instance-independent theoretical guarantees on prediction error are available for such out-of-sample tests, motivating a purely empirical backtest [41]. This is the market-microstructure statement of the performativity problem: deployed models recursively influence the data distributions on which they are evaluated through agents' strategic responses [15], and what a decision-maker would want is an equilibrium in which the model is optimal for the distribution it induces, the stable points of retraining [17].
Related but non-identifying evidence. The corpus cites, without reporting results, work indicating that execution quality and dealer behavior are altered by electronic RFQ trading (O'Hara et al.) and that bilateral relationships materially affect trading costs in corporate bonds (Jurkatis et al.), concluding that quoting models should combine inventory-sensitive pricing with a realistic description of how RFQs translate into fills [6]. It also points to literatures on competition and learning in dealer markets and on algorithmic collusion in MD2C platforms [20][22][24], and to explicit strategic formulations (Stackelberg/Nash) for competitive liquidity provision [23]. These are reference-level citations only.
Established by the corpus: (a) the mechanism is a short-window, multi-dealer, best-respondent auction with censored feedback [1][2][4][32]; (b) the policy is a multi-dimensional ladder of offsets by bond, tier, side and size, jointly with inventory/hedging and client-targeting choices [34][5][21][30]; (c) hit ratio, ranking/cover, and multi-horizon flow value are the observable outcomes, and hit ratio is used by clients to route future RFQs [34][4][29][21]; (d) the mapping from quotes to outcomes must be treated causally because dealers condition their quotes on the same variables that drive client behavior [7][10][30]; and (e) there are at least two feedback loops, within-RFQ selection/winner's curse [23] and cross-RFQ score-gated routing [24][8], that make flow, win rate, adverse selection and realized spreads functions of the dealer's own past policy [1][8].
Not established by the corpus: any quantified causal estimate, no reported elasticity of hit ratio to spread, no measured change in invited-panel probability following a change in quoting aggressiveness, no experimental or event-study magnitudes for adverse selection versus panel size in corporate bonds, and no empirical validation of the bifurcation/hysteresis prediction in [8]. The corpus's own prescription for closing that gap is either back-door identification on a sufficient conditioning set from dealer histories [7][27] or randomized controlled trials / A/B tests where identification from history fails [40].
Performative prediction replaces the classical assumption of a fixed data-generating process with an explicit map from deployed parameters to data distributions. The central mathematical object is the distribution map from the parameter space into the space of probability measures over data, so that deploying a model with parameters θ causes the data subsequently observed to be drawn from the induced distribution D(θ); the performative risk of θ under a loss ℓ is then the expected loss evaluated on D(θ) itself [14]. This formalizes the recursive feedback in which "the act of deploying a predictive model shifts the data distribution itself, often as a result of the modeled agents' strategic, behavioral, or otherwise performative responses to predictions" [15].
The motivating diagnosis is that such feedback is pervasive and, when ignored, masquerades as distribution shift: a bank that predicts elevated default risk and responds with a high interest rate further increases the customer's default risk, so that "the bank's predictive model is not calibrated to the outcomes that manifest from acting on the model"; analogous loops arise in traffic prediction, crime-location prediction, recommendation, and stock-price prediction driving trading activity and hence prices [17]. Because the decision-maker acts according to the model, "the distribution over data points appears to change over time," and the standard practical response, frequent retraining, is usually framed as "an undesired, yet necessary, cat and mouse game of chasing a moving target" [17].
Two structural specializations in the corpus are worth noting for a dealer-quoting application. First, strategic classification is exactly a performative prediction problem with a Stackelberg structure, "since agents adapt their features only after the bank has deployed their classifier," and the induced distribution is generated by problem-specific best-response functions applied to a baseline distribution [25]. Second, reinforcement learning is itself a case of performative prediction: "the choice of policy … affects the distribution over … the set of visited states … and actions … in a Markov Decision Process" [16].
Stability (fixed point). Performative stability is a fixed-point condition: a parameter is performatively stable when it "is optimal with respect to its own induced distribution" [14]. Perdomo et al. state it via the decoupled performative risk DPR(θ, θ′), the expected loss of θ′ on the distribution induced by θ, with stability requiring that θ minimize DPR(θ, ·) [11]. The interpretive content is precisely the elimination of chasing: "A performatively stable model minimizes the expected loss on the distribution resulting from deploying [it] in the first place. Therefore, a model that is performatively stable eliminates the need for retraining after deployment since any retraining procedure would simply return the same model parameters. Performatively stable models are fixed points of risk minimization" [11]; equivalently, "performatively stable models … achieve minimal risk for the distribution they induce and hence eliminate the need for retraining" [13]. In algorithmic terms, a performatively stable point is "a fixed point of the RRM operator" [12], and "the fixed points of retraining are performative stable points" [16].
Optimality. The performative optimum is the parameter that directly, globally minimizes the performative risk PR(θ) [14], i.e. it "globally minimizes … but is not necessarily a stable point" [12].
The two concepts are distinct. "Performative optimality and performative stability are in general two distinct solution concepts. Performatively optimal models need not be performatively stable and performatively stable models need not be performatively optimal" [11]. The corpus is explicit that this gap is not merely technical: a performatively stable point "may be suboptimal in terms of performative risk," "the gap between stability and optimality is generally nonzero," and "this gap is central to both theory and the practical pathology of performative settings" [12]. Consequently, a natural theoretical question, and one that Perdomo et al. pose directly, is "under what conditions can we find models that approximately satisfy both?" [13]; the paper establishes "properties of the objective under which stable points and performative optima are close" [13].
Equilibrium interpretation. For strategic classification, the institution's optimal strategy is the Stackelberg equilibrium, "the classifier which achieves minimal loss over the induced distribution in which agents have strategically adapted their features", and this "equilibrium notion exactly matches our definition of performative optimality" [25]. So in that setting the optimum, not the stable point, is the Stackelberg solution, while repeated retraining corresponds to RRM and converges (under conditions) to the stable point [25].
Robust variants. A recent extension is the distributionally robust performative optimum (DRPO), which "address[es] model misspecification by minimizing worst-case risk over a divergence ball around the nominal distribution map" [14], relevant when the analyst cannot commit to a single estimate of how the market responds.
The corpus identifies three ingredients that jointly control existence, uniqueness, and convergence: (i) sensitivity of the distribution map, (ii) strong convexity of the loss in the parameters, and (iii) smoothness.
- Sensitivity / Lipschitzness of the distribution map. Perdomo et al.'s "algorithmic analysis of these methods reveals the existence of stable points under the assumption that the distribution map is sufficiently Lipschitz," and they "identify necessary and sufficient conditions for convergence to a performatively stable point" [13]. The summary source states the sensitivity condition explicitly as Lipschitz sensitivity of D(·) in the parameter, "e.g., in the Wasserstein-1 or [total variation] topology" [12].
- Strong convexity and smoothness. "Unique stable points" are obtained "under the assumption that the objective is strongly convex and smooth" [19]. Notably, weak convexity is insufficient: "(weak) convexity alone is not enough. Performativity thus gives another intriguing perspective on why strong convexity is desirable in supervised learning" [16].
- A composite contraction condition. The three conditions, strong convexity, smoothness, and Lipschitz sensitivity, jointly deliver linear convergence to the unique performatively stable point [12]. The contraction rate is governed by a stability parameter combining the sensitivity of the distribution map, joint smoothness, and strong convexity; "when [this quantity is below one], RRM converges rapidly" [12]. (The corpus renders these formulas with symbols stripped, so the exact algebraic form of the threshold and rate is not recoverable from the evidence here, only the qualitative structure "sensitivity × smoothness / strong convexity" [12].)
- Existence without strong convexity, via constraints. Existence of stable points can be obtained "under weaker assumptions on the loss, in the case where the solution space is constrained" [19], noteworthy because, in contrast, performative optima "are always guaranteed to exist" (over the extended real line), whereas "it is not clear whether performatively stable points exist in all settings" [19].
- Hardness beyond the contraction regime. When the stability parameter exceeds one, i.e. "strong performative effects or weak regularity", "fixed-point computation becomes computationally hard (PPAD-complete …), even in quadratic and linear settings" [12]. This is a sharp negative result: the fixed point may still exist as an object, but no efficient algorithm should be expected in general.
Non-uniqueness and hysteresis in a market analogue. The corpus contains no performative-prediction result on multiplicity of fixed points beyond the hardness statement above. However, a closely related feedback structure in aggregator-routed RFQ market making, where routing intensity is multiplied by a promotion gate driven by the dealer's slowly varying win score, and where trade outcomes update that score, is shown to exhibit exactly the pathologies one would expect outside a contraction regime: "For steep (logistic) promotion gates, the score dynamics can exhibit fold bifurcations, bistability, and hysteresis, producing an endogenous 'campaign vs. harvest' pattern in optimal quoting," confirmed numerically [8]. That paper obtains this via an explicit timescale separation, "an adiabatic approximation that separates fast inventory dynamics from slow score dynamics," yielding a one-dimensional score drift field [8], and it identifies background (non-gated) flow as a stabilizing force [8]. This is suggestive rather than a theorem about performative stability: it comes from a stochastic-control model of dealer quoting, not from the performative-prediction literature, and the mapping between the two formalisms is not established in the corpus.
Repeated risk minimization (RRM). RRM is "retraining, formally referred to as repeated risk minimization …, where the exact minimizer is repeatedly computed on the distribution induced by the previous model parameters" [13]; it is realized in practice as "models are iteratively retrained on data generated by the distribution induced at the previous step" [14]. The key conceptual reframing is that retraining is not a nuisance but "the natural equilibrating dynamic for performative prediction," whose fixed points are performatively stable points, and which "converges to such stable points under natural assumptions, including strong convexity of the loss function" [16]; equivalently, the desired equilibrium, "a certain equilibrium where the model is optimal for the distribution it induces", coincides "with the stable points of retraining, that is, models invariant under retraining" [17]. Quantitatively, RRM converges to a unique performatively stable point "if the sensitivity parameter is small enough" [18], and converges linearly under strong convexity + smoothness + Lipschitz sensitivity [12].
Repeated gradient descent (RGD). RRM's drawback is that it "requires access to an exact optimization oracle" [18]. RGD relaxes this: parameters "are incrementally updated using a single gradient descent step on the objective defined by the previous iterate," and is introduced "as a computationally efficient approximation of RRM, which … adopts many favorable properties of RRM" [13]. Formally, "a simple gradient descent algorithm also converges to a unique stable point" [18].
A convergence-region gap between the two. Perdomo et al. flag an asymmetry: as the step size of repeated gradient descent tends to a limit, "this procedure converges for [a restricted range of sensitivity]," whereas "exact repeated risk minimization … provably converges for every [sensitivity below the stated threshold], and we showed this inequality is tight"; whether the gap "is a fundamental difference between both procedures or an artifact of our analysis" is left open [16]. (Again the numerical thresholds are not rendered in the available text [16].)
Population versus finite sample. The analysis is carried out "at a population level" and then extended "to finite samples" [13].
Two-timescale stochastic approximation and derivative-free/zeroth-order methods. The corpus contains no evidence on performative-prediction algorithms of these types, no two-timescale stochastic approximation scheme for jointly updating a policy and an estimated response model, and no derivative-free/zeroth-order or gradient-based performative-optimum-seeking procedures (as opposed to stability-seeking RRM/RGD). This is a genuine gap in the evidence available here, not a claim that such methods do not exist. The only timescale-separation argument in the corpus is the adiabatic (fast-inventory / slow-score) reduction in the RFQ market-making model [8], which is a modelling approximation for HJB analysis rather than a convergence guarantee for a learning algorithm.
A distinct and important limitation is that the classical performative-prediction results define outcomes at the population level. Existing work "almost exclusively defines 'actual outcomes' at the population level, focusing on stability and optimality under repeated risk minimization while evading the question of whether one can generalize from a finite sample to the population" [35]. Once "predictions react back upon the data distribution, standard conclusions regarding generalization from training to test sets no longer apply" [35], and there is an inherent tension: "the more a model is used to intervene in the data …, the more the sample deviates from the original population, paradoxically making it harder to reliably infer population properties" [35].
Recent learning-theoretic work addresses this without assuming a functional form for the response map, requiring only Wasserstein sensitivity, and "incorporate[s] performative drift into generalization bounds using covering numbers and Wasserstein distances" [36]. The proof decomposes three Wasserstein segments, sample-to-population convergence, in-sample performative drift, and population performative drift, converts them into expected-loss differences via Kantorovich–Rubinstein duality, and (because drift can push evaluation points outside the support of the base distribution) measures hypothesis-class richness with a covering-number entropy integral rather than Rademacher complexity; a key observable driving the bound is the "performative response rate m/n," the fraction of sampled units whose behaviour changed due to the prediction [37]. Two failure directions are characterized as optimization problems in Wasserstein space: population self-negation, whose worst case is an inf-sup functional matching Distributionally Robust Optimization over a Wasserstein ball, and sample self-fulfilment, where retraining-based empirical risk minimization is equivalent to an inf-inf functional corresponding to Distributionally Favorable Optimization [37]. In other words, a naively retrained policy can converge to an empirical echo chamber, a fixed point of its own induced sample rather than of the true response map [37].
Two statements in the corpus sit in productive tension and should be read together rather than as a contradiction. Stability is presented as a desideratum: a stable model "eliminates the need for retraining after deployment" because retraining "would simply return the same model parameters" [11], and it "achieve[s] minimal risk for the distribution [it] induce[s]" [13]. Yet the same fixed point "may be suboptimal in terms of performative risk," and the stability–optimality gap "is generally nonzero" [12]. For a quoting policy, then, converging to a fixed point is a self-consistency guarantee, not a profit-maximality guarantee, and the corpus offers only the existence of conditions "under which stable points and performative optima are close" [13], without those conditions being spelled out in the retrievable text.
The starting point of the OTC-dealer literature represented in this corpus is that execution is not an exogenous arrival process. In multi-dealer-to-client (MD2C) request-for-quote (RFQ) markets, a client solicits a small panel of dealers and trades, if at all, with the best respondent, which makes market making "a joint problem of inventory management, counterparty selection, and endogenous execution probability: a dealer must quote competitively enough to win flow, but not so aggressively as to destroy margin or accumulate undesirable inventory" [1]. Execution is "shaped not only by price, but by the RFQ protocol itself, the number of requested dealers, the response behavior of competitors, and the search technology available to clients" [1]. The synthesis drawn explicitly in that survey is the key equilibrium claim: "the dealer's hit ratio is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1][5].
Structurally, the mechanism that generates this endogeneity is well documented. Clients request prices from multiple liquidity providers and either trade with one or do not trade at all [5]; on MD2C platforms the process is only partially observable, the winner learns the second-best "cover" price, others learn their ranking and (sometimes) whether a trade occurred at all [4]. This "fosters competition but creates a highly strategic environment for pricing, in which dealers must optimize quotes under uncertainty about competitors' prices and client preferences" [4]. The dealer's objective in this environment is explicitly a trade-off: "balancing the probability of winning a trade with expected profitability and inventory risk," where aggressive quoting raises hit probability but compresses margin and exposes the dealer to adverse selection and post-trade price moves [3][4].
The corpus documents three distinct modelling strategies for the joint quoting/flow problem.
(i) Structural estimation of the primitives that generate the hit-ratio function. Fermanian, Guéant and Pu provide "a foundational empirical model of European corporate-bond RFQs, estimating dealer-quote and client-reservation distributions and explicitly deriving ex-ante hit-ratio functions" [1][22]. This is the cleanest available statement of the joint-equilibrium logic: the dealer's win probability is derived from the distribution of rival quotes and the distribution of client reservation prices, rather than assumed. Related structural work makes hit probabilities depend on the competitive set and on the number of solicited dealers, so that "how many dealers are asked" is described as "a first-order primitive in RFQ economics" [5].
(ii) Search-and-auction and search-frictions models. The corpus records that Hendershott and Madhavan analyse the trade-off between search and auction in OTC markets, that Kargar et al. study sequential search in corporate bonds, and that Wang shows multi-dealer platforms "face structural limits in how much price competition they can generate" [1]. The corpus does not report the internal structure, equilibrium conditions, or quantitative predictions of these search-and-bargaining models, only their existence and headline theme, so any statement about, e.g., bargaining-power splits or intermediation-chain pricing in that tradition would be unsupported here.
(iii) Reduced-form competition inside a control problem. A large part of the market-making literature does not solve for rival strategies at all: stochastic optimal control approaches combine inventory management with quote-dependent order arrivals [5], and one strand "treats competition in reduced form: a reference market maker optimises quotes while fill probabilities depend on differences versus exogenous competing quotes, yielding approximate closed-form solutions under linear-quadratic objectives" [23]. Extensions add size heterogeneity via multi-size ladder controls and non-local (jump) HJB operators indexed by trade size [5], hedging and market impact, capturing "the practical transition between pure internalization and external risk transfer" [21], and Markov-modulated order-flow models of RFQ liquidity dynamics [21].
There is an explicit tension worth flagging between (i)/(iii) and practice. Although the hit ratio is treated as an equilibrium object [1], "most of this literature treats execution probability in reduced form, typically through quote-dependent intensities or fill probabilities, rather than elevating hit ratio itself to an explicit control target" [21], even though "in industry practice, historical hit ratios are used to select dealers and route RFQs, especially for liquid electronic bond orders, alongside axes, response quality, and relationship information" [21]. One response is to make the KPI an explicit objective: quote offsets are chosen jointly with a penalty for deviations from a desired hit-ratio objective, with fill probabilities (f) mapping offsets to OTC fill intensities, a size-weighted instantaneous hit ratio per bond/tier, target levels per targeted client tier, and hit-ratio weights entering the dealer's optimisation alongside running inventory-risk and terminal-inventory penalties [21][34]. This is a control-theoretic form of self-consistency (the policy is consistent with a chosen target), not an equilibrium one, and [1] is precisely the caution that the target itself is not a free parameter, since it is co-determined by rivals, platform design and client search.
The adverse-selection channel in this corpus is a selection-on-winning mechanism rather than a sequential-trade Bayesian-updating mechanism. Oomen argues that the "winner's-curse" in an aggregator "can strengthen as more LPs compete: the best displayed quote is more likely to be selected precisely in states that are unfavorable to the winning LP," which provides "an economic rationale for limiting the number of LPs exposed to each RFQ and for using routing and ranking rules rather than always inviting the full pool" [23][1]. In control form, this appears as an explicit correction: "conditional on winning, trades earn spread minus an adverse selection correction and contribute to inventory risk" [8]. Informational risks accompanying competitive quoting, and internalisation-versus-externalisation of client flow, are cited as further motivations for jointly modelling client access, dealer incentives and platform design [23]; a dedicated strand studies optimal quoting under adverse selection and price reading [20][9].
The joint prediction that emerges is that a self-consistent quoting policy is characterised by an interior optimal win probability: pushing win probability up mechanically worsens both the margin and the conditional quality of won flow [3][4][23], while pushing it down forgoes inventory-mixing capacity [8]. The two-tier model makes this mapping explicit, using an envelope-theorem argument to "express optimal controls through derivatives of the one-dimensional reduced Hamiltonians, yielding an interpretable mapping from optimal win probabilities to optimal offsets" [8]. Note that a Glosten–Milgrom-style sequential-trade model of spreads as compensation for the probability of informed trade is not present in this corpus; the adverse-selection results cited here are auction/aggregator selection effects and reduced-form corrections [8][23].
The sharpest available result on self-consistent quoting under endogenous flow concerns platform routing. Two empirical facts are highlighted: "(i) the competitive set is typically a subset of all eligible LPs and is often shaped by ranking, relationship, or platform-level 'top list + exploration' logic; (ii) inclusion and rank depend on slowly varying performance scores (e.g., long-run win ratio, response quality), which in turn depend on how aggressively the LP quotes" [24]. The corpus also notes that, while RFQ micro-competition and aggregator economics are separately well studied, "there is limited academic control literature that treats platform shortlisting driven by a slow win-score as an explicit state variable feeding back into the LP's future opportunity intensity" [24].
Modelling that loop explicitly, an aggregator tier whose opportunity intensity is multiplied by a promotion gate driven by the dealer's win score and whose outcomes update the score, plus a non-gated background tier that does not update the score [8][24], yields a two-scale problem with an adiabatic separation of fast inventory dynamics from slow score dynamics, a quadratic inventory ansatz, a quasi-stationarity inventory-curvature scaling and a one-dimensional score drift field [8]. The central prediction is that self-consistency can fail to be unique: "For steep (logistic) promotion gates, the score dynamics can exhibit fold bifurcations, bistability, and hysteresis, producing an endogenous 'campaign vs. harvest' pattern in optimal quoting," with numerical experiments confirming this and highlighting "the stabilizing role of background flow in maintaining inventory-mixing capacity even when the dealer is weakly promoted" [8]. In other words, the equilibrium of quoting and flow may support multiple locally self-consistent regimes, invest in score by quoting tight, or harvest margin while lightly promoted, with path dependence between them [8].
Beyond reduced-form competition, the corpus identifies two further strands. First, "other works move to explicit strategic settings (e.g., Stackelberg or Nash structures), providing equilibrium characterisations for competitive liquidity provision" [23][24], cited instances include market making with competition as a Stackelberg equilibrium, Nash equilibrium between brokers and traders, and liquidity competition between brokers and an informed trader [22]. Second, "a parallel strand uses reinforcement learning and agent-based models to study the emergence of tacit coordination in electronic markets and in MD2C-type dealer platforms, and the role of heterogeneity in mitigating such effects" [23][24], including work on algorithms and supracompetitive prices with tick-size effects, AI-driven liquidity provision in OTC markets, competition and learning in dealer markets, and algorithmic collusion in MD2C platforms [22].
What these add relative to single-agent control is therefore threefold, on the corpus's own terms: (a) equilibrium characterisation rather than best response to an exogenous competitor distribution [23]; (b) hierarchy, Stackelberg structures capture a leader who commits to a policy anticipating followers' adaptation [23][25]; and (c) emergent multi-agent phenomena that no single-agent objective encodes, notably tacit coordination and supracompetitive pricing among learning dealers, and the mitigating role of heterogeneity [23][24][22].
Evidence gap: mean-field games. The corpus contains no material on mean-field games, mean-field control, or mean-field equilibria in dealer markets. The closest analogues present are the reduced-form treatment of competition through an exogenous distribution of competing quotes [23] and the aggregate-score/routing feedback of the two-tier model [8][24]; any claim about mean-field limits, master equations, or many-dealer aggregation would go beyond this evidence.
Performative prediction formalises the case where "deploying a predictive model shifts the data distribution itself, often as a result of the modeled agents' strategic, behavioral, or otherwise performative responses to predictions" [15]. Its two solution concepts are precise. Performative stability is the fixed-point condition that a parameter "is optimal with respect to its own induced distribution" [14][11]; such a model "eliminates the need for retraining after deployment since any retraining procedure would simply return the same model parameters," and stable points are "fixed points of risk minimization" [11][13]. Performative optimality globally minimises performative risk and "is not necessarily a stable point" [12][14]. Critically, the two differ: "performatively stable points… may be suboptimal in terms of performative risk," and "the gap between stability and optimality is generally nonzero" [12].
Three differences from OTC-market equilibrium concepts follow directly.
-
Stability is single-agent best response to one's own footprint; Nash/Stackelberg equilibria are joint conditions. A performatively stable quoting policy is one that a dealer would not revise after observing the flow it induces [11][14]. But in RFQ markets the induced flow depends on "rival participation" and on platform routing [1][24], so stability with respect to one's own induced distribution is strictly weaker than an equilibrium in which rivals are simultaneously best-responding, as delivered by explicit Nash or Stackelberg formulations [23][22]. The corpus supports this distinction but does not contain a formal theorem relating dealer-market Nash equilibria to performative stability, so the mapping should be read as conceptual.
-
Stackelberg equilibrium corresponds to performative optimality, not stability. In strategic classification, "agents adapt their features only after the bank has deployed their classifier," and the Stackelberg equilibrium, "the classifier which achieves minimal loss over the induced distribution in which agents have strategically adapted", "exactly matches our definition of performative optimality" [25]. Repeated retraining on induced distributions is exactly repeated risk minimisation, whose fixed points are stable rather than optimal points [25][16][12]. Translated to dealing: a dealer who repeatedly recalibrates a hit-probability model on its own realised flow is running an RRM-type dynamic that converges (when it converges) to a stable, possibly sub-optimal, quoting policy [16][12], it does not by construction internalise the leader-side value of moving win probability to change future routing, which is precisely the effect the promotion-gate model makes first-order [8][24].
-
Convergence conditions map onto the bistability result. RRM and related iterations "converge linearly to the unique performatively stable point" under strong convexity, smoothness, and Lipschitz sensitivity of the distribution map, with the contraction governed by a stability parameter; when performative effects are strong or regularity weak, "fixed-point computation becomes computationally hard (PPAD-complete…), even in quadratic and linear settings" [12]; stable points exist "under the assumption that the distribution map is sufficiently Lipschitz" [13][19], and retraining converges under strong convexity, weak convexity alone is not enough [16][18]. The dealer-market counterpart of a high-sensitivity distrib
In the RFQ setting the "market-response function" is, operationally, the hit (fill) probability: the probability that a client trades with the dealer given the quoted spread and the context of the request [10]. This object is what closes the loop between pricing and P&L: the dealer's optimal spread solves an implicit first-order condition in which the derivative of the hit probability with respect to the spread appears, so any bias in the response function maps directly into mispricing [10]. The economic reason the identification problem is hard is that this probability is not a primitive of the environment. The dealer's hit ratio "is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], and execution depends "not only by price, but by the RFQ protocol itself, the number of requested dealers, the response behavior of competitors, and the search technology available to clients" [1]. Structural work on multi-dealer-to-client (MD2C) platforms makes the same point quantitatively: hit probabilities depend on the competitive set, so "how many dealers are asked" is a first-order primitive of RFQ economics [5].
The core confounding mechanism is that dealers do not quote on exogenous, context-independent grounds. They "actively incorporate information from the market environment, the RfQ itself, the characteristics of the instrument, the identity of the client, and the competitive landscape, among others, when setting prices," which "introduces correlations between historical spreads and RfQ outcomes that may not reflect causal relationships" and are therefore irrelevant, or actively misleading, for choosing revenue-maximising spreads [7]. In the graphical formalisation, client identity is known to quoting dealers on MD2C platforms, so pricing policies absorb client features, demand and price-sensitivity drivers, bond features such as liquidity and relative value, and RFQ features such as the number of dealers in competition; if the dealer holds an axe, prices are skewed to trade [30]. These variables therefore sit on back-door paths and "introduce spurious dependencies between the quoted price and the RfQ outcome", they are confounders in the causal-inference sense [30]. Pricing policies are further shaped by client segmentation or product-specific rules, so "the historical relationship between quotes and outcomes may not reflect the causal effect of intervening on prices" [3].
The RFQ protocol itself censors outcome information asymmetrically. On MD2C platforms the winning dealer is told the second-best quote (the "cover price"), other dealers receive only their ranking and whether a trade occurred, and in some cases losing dealers cannot even distinguish a missed RFQ (traded away) from a passed RFQ (no trade at all) [4]. Two consequences follow. First, the competitor-quote distribution, the object that actually determines the win/lose boundary, is observed only through a coarse, win-conditional channel [4]. Second, any revenue-based target is observed only on won trades, whereas the response function must be estimated over all requests.
The corpus documents one concrete mitigation: framing estimation as a binary classification of won versus lost (missed or passed), which "enables estimation of trading probability without conditioning on observing actual transactions" [33], i.e. keeping the unwon RFQs in the estimation sample rather than selecting on execution. A second, distinct problem is that latent variables can bias the model even after observable confounders are adjusted for: neglecting "partially observable factors, including information asymmetry or client intent (e.g., price discovery)" can bias predictive models and lead to suboptimal decisions [27]. In the empirical exercise of that same framework, information asymmetry is simply assumed away (set to zero), an assumption acknowledged as potentially violated in practice and defended only as a common structural bias across the models being benchmarked [33].
There is also a measurement-side selection issue in how "response" is scored. Because round-trips in illiquid bonds span long horizons, revenues are often marked at a horizon using market mid-prices, end-of-day flow value, sometimes with liquidity penalties for unrealised mark-to-market profit, while short-horizon flow value (seconds to minutes) is used precisely because it makes information-asymmetry-driven losses easier to detect before market factors dominate [29]. The choice of horizon therefore determines whether adverse selection is attributed to the client or absorbed into noise [29].
The corpus supports a fairly specific programme:
- Interventional, not conditional, targets. The quantity of interest is the distribution of outcomes under an intervention on the spread (do-operator), not the historical conditional distribution; this "requires a causal, rather than purely associative, relationship between the variables" [10]. Mid-price is treated as exogenous and agreed between client and dealer, and is factored out [10].
- Back-door adjustment with a minimal conditioning set. Causal inference, specifically the back-door criterion, identifies the minimal set of variables required for valid causal effect estimation [7]. For the hit-probability model this set is volatility plus RFQ features, bond features and client features [7]. Importantly, including all available features indiscriminately "can introduce spurious correlations, distorting the estimation of causal effects" [10]. If the conditioning set is smaller than required, the criterion can still be applied by integrating over the distribution of the omitted variables [7]. Note also that a minimal conditioning set does not force the dealer into price segmentation, although segmentation can raise revenues when feasible [7].
- Graphical models to expose latent structure. Encoding the RFQ workflow as a probabilistic graphical model makes observed and latent dependencies explicit and supports counterfactual reasoning about interventions such as setting quotes or initiating client outreach [3]; the published framework applies this to optimal pricing, revenue-potential estimation under fixed pricing policies, and axe/client matching (an uplift-style intervention) [3][27].
- Experimentation when identification fails. Where do-calculus shows the effect is not identifiable from historical data, "it becomes necessary to design experiments, specifically, randomized controlled trials (RCTs), commonly implemented as A/B tests, to empirically measure the impact of the intervention" [40].
- Estimation machinery. One may specify a full generative model of the joint distribution and estimate by MLE/MAP/Bayes, or fit discriminative models for each required conditional [38]. Discriminative models "often outperform generative models in predictive tasks" and benefit from regularisation and ensembling for out-of-sample generalisation, but are more opaque about structural mechanisms [38]. In practice, feature selection is guided by the back-door minimal set combined with domain knowledge and Random-Forest importance [33]. Related finance work uses Double Machine Learning to estimate structural relationships more reliably, and causality-inspired forecasting for robustness under distribution shift [26].
- Counterfactual pricing on MD2C platforms is now an established use of causal interventions, "bridging structural RFQ models with modern discriminative learning" [5].
Evidence gaps. The corpus contains no material on classical off-policy evaluation estimators (inverse propensity scoring, doubly robust estimation, off-policy policy-gradient), no instrumental-variable or propensity-score design specific to RFQ quoting, and no survival/censored-regression treatment of unwon RFQs beyond the won/lost relabelling in [33]. It also does not contain an identification strategy for a price-impact or liquidity function as such; the market-impact material is model-side (hedging and market impact embedded in dealer-market control models [21], the observation that impact and price jumps often appear without association to prior events [32]), not an observational identification result.
Even a correctly de-confounded static hit-probability model leaves out a second-order channel that is documented in the corpus: platform routing itself responds to quoting behaviour. In aggregator-routed RFQ markets the competitive set is typically a subset of eligible providers, shaped by ranking, relationship or "top list + exploration" logic, and inclusion and rank depend on slowly varying performance scores such as long-run win ratio and response quality, "which in turn depend on how aggressively the LP quotes" [24]. Modelling this explicitly, a promotion gate driven by the dealer's win score, with score-updating and non-updating flow tiers, can produce fold bifurcations, bistability and hysteresis, and an endogenous "campaign vs. harvest" pattern in optimal quoting [8]. This is precisely the structure that performative prediction formalises: deployment shifts the distribution the model is later evaluated on, and the natural equilibrium concept becomes a fixed point at which a model is optimal for the distribution it induces [15][16]; retraining is then the equilibrating dynamic rather than a nuisance, converging to such stable points under strong convexity and sufficient-Lipschitz distribution maps [16][13]. The evidence does not, however, show anyone applying performative-prediction machinery to RFQ pricing; that connection is an inference across two literatures in the corpus rather than a documented result.
RFQ is "the most widely used electronic execution protocol in bond markets" [2] and remains the dominant electronification path in corporate bonds [1]. On the buy-side/platform layer, execution management systems already automate the mechanics: "responding to RFQs, classifying trades, routing orders, can be automated within the EMS," alongside liquidity aggregation, protocol selection, pre-trade price comparison and post-trade TCA used to demonstrate best execution under MiFID II [47][48][49].
On the dealer side, the documented practical setting is a blind auction: inquiries arrive over automated trading systems as RFQs, with "subsequent algorithmic responses with a quote, conditional on the dealer's inventory interest and trading strategy," in a market driven by institutional flow and characterised by sparse pricing data and high dimensionality of inquiry characteristics and market conditions; "a key problem in this business is the execution likelihood estimation or fill probability of a given RFQ response" [28]. The described architecture decouples market-state representation from the learned state-to-outcome mapping: an offline component generates discriminative features, and an online component either infers fill probability for incoming inquiries or recalibrates to new state-to-outcome associations [32]. Practical constraints are explicit: model complexity considerations in model risk management and regulatory controls favour interventions that change only input data rather than the learning algorithm, so that a less complex or more robust model can be retained [50]. Industry usage of the response object is likewise documented: historical hit ratios are used to select dealers and route RFQs for liquid electronic bond orders, alongside axes, response quality and relationship information [21].
The evidence points mainly to the latter, with two partial exceptions.
Treated as exogenous. The clearest statement comes from the fill-probability backtesting literature: any performance gain from embedding a new fill-probability estimator "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs"; the response is therefore acknowledged and then deliberately excluded by "keeping the test environment fixed" via backtesting [39]. The associated methodology is explicit that with strict temporal decoupling of train and test sets, "no general and data-instance independent theoretical guarantees on prediction errors can be made," so the approach is "purely empirical" [41]. On the modelling side, most of the OTC market-making literature "treats execution probability in reduced form, typically through quote-dependent intensities or fill probabilities, rather than elevating hit ratio itself to an explicit control target" [21].
Partial exceptions. First, the causal MD2C framework does target the interventional response: it estimates the hit probability under an intervention on the spread, adjusts for the confounders that generate the endogeneity, and reports that ignoring confounders or partial observability biases predictive models and yields suboptimal decisions [10][7][27]. That is an explicit model of the response to the dealer's own action, but within a static, single-RFQ setting; the corpus reports empirical benchmarking of generative versus discriminative implementations [38][33] and does not establish live production deployment. Second, control-theoretic work makes the induced response a state variable: hit-ratio targeting is written into the dealer's objective, with size-ladder quotes, tier-specific hit-ratio targets and inventory penalties [34], and the two-tier win-score model makes future opportunity intensity depend on past quoting through the platform's promotion gate [8][24]. These are model-based and validated numerically [8]; no deployment evidence is in the corpus.
Conflict to flag. [39] and [8]/[3] represent two positions in the corpus: the empirical fill-probability pipeline treats induced market response as unmeasurable and holds the environment fixed for benchmarking [39], whereas the causal and stochastic-control strands argue the response to the dealer's own quotes is exactly what must be modelled [10][8]. Both are cited above; the corpus does not adjudicate them with head-to-head evidence.
If quoting policies are retrained on outcomes generated by their own prior quotes without treating the induced shift explicitly, the training distribution is no longer the population of interest. The corpus's most direct evidence on degenerate self-referential training is from generative AI, where indiscriminate learning from model-produced data causes "model collapse", progressive loss of the true distribution's tails and convergence to low-variance point estimates, attributed to functional approximation, sampling and learning errors [44][43][46]. That literature concerns generative models trained on synthetic data, not dealer quoting, so it is an analogy rather than evidence about RFQ systems; the transferable point is only that closed training loops require deliberate design. The mechanisms the corpus does document as directly relevant are the win-score feedback loop with hysteresis [8][24] and the performative-stability fixed-point framing [15][16], plus the recommendation to fall back on randomised experiments where observational identification fails [40].
Corporate-bond RFQ pricing is not a one-shot prediction problem. Beyond the direct ("first-order") performative link in which a dealer's own quote changes the probability that the trade is won, the object the dealer must estimate interventionally rather than associatively [10], there is a slower, second-order channel: quote and trade decisions leave traces in the shared measurement infrastructure (composite/evaluated price surfaces, TCA and best-execution reporting, hit-ratio and response-quality scores, regulatory trade reporting), and that infrastructure is precisely what supplies the features, benchmarks, and labels of the next generation of the pricing model. This section traces that channel and then examines how market impact and inventory dynamics make the response endogenous over time in corporate bonds specifically.
The relevant institutional fact is that the dominant execution protocol is RFQ [2], in which a client solicits a panel of dealers and typically trades, if at all, with the best respondent [1]. Feedback to participants is deliberately partial: the winner learns the second-best ("cover") quote, other dealers learn only their ranking and whether a trade occurred, and in some cases cannot even distinguish a missed from a passed RFQ [4]. So the raw label-generating process is already a censored function of the competitive interaction, not an exogenous draw.
Around this protocol sits a pricing and analytics layer built from the same quotes and trades. Execution management systems aggregate dealer pricing and multi-dealer-platform liquidity into a single real-time view, present side-by-side price comparisons "with live market data to benchmark incoming quotes," and then run post-trade transaction cost analysis that measures performance against benchmarks and characterises dealer relationships [47]. That analytics output is not decorative: MiFID II requires firms to take sufficient steps to achieve best execution and to retain evidence, which the EMS supplies via automated pre-trade comparisons, a captured snapshot of market conditions at execution, and post-trade TCA [48], with audit-trail obligations explicitly cited as a driver of electronification [49]. On the regulatory side, FINRA has responded to increased trading of TRACE-eligible securities in dark pools and "greater use of request-for-quote processes" by requiring firms to report additional information, including the ATS's unique MPID, to TRACE [51], i.e., the reporting layer is being expanded to capture exactly the venue/protocol dimension along which quoting behaviour varies.
The contamination becomes concrete when one looks at what a dealer model actually ingests. In the MD2C causal-modelling framework, the central pricing feature is the normalized half-spread, the dealer's quoted spread "adjusted by a market liquidity proxy, such as half the bid-ask spread of a composite benchmark like CBBT", alongside yield volatility computed as the standard deviation of daily bond-yield changes over the past 30 days [33]. On the label side, revenue is measured against market mid-prices: an end-of-day flow value that uses market mids to value inventory, sometimes with explicit liquidity penalties on unrealised mark-to-market profit, and a short-term flow value measured against the mid a few seconds or minutes after the trade [29]. Both the treatment variable (a spread normalised by a composite) and the reward label (P&L marked to a mid) are therefore functions of the observable price surface, not of primitive quantities.
An important limit of the corpus must be stated plainly: nothing here documents the mechanical construction of composite or evaluated prices, or of TRACE-derived spread statistics, from contributed dealer quotes and reported trades. The evidence establishes (i) that dealer features and labels are defined relative to composite mids and bid-ask benchmarks [33][29] and (ii) that quotes and trades are aggregated and published into pre-trade comparison views, TCA, and TRACE reporting [47][48][51]; the closing step, that a given dealer's own quoting materially moves those composites and hence its own future features and labels, is a plausible inference but is not directly evidenced in this corpus. It should be treated as a hypothesis to be measured, not an established magnitude.
Where the corpus is unambiguous is on a related and arguably stronger second-order channel: performance statistics computed from a dealer's own quoting decisions are used to decide which future RFQs it will even see. In industry practice, "historical hit ratios are used to select dealers and route RFQs, especially for liquid electronic bond orders, alongside axes, response quality, and relationship information" [21]. Consistent with this, the hit ratio "is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], and structural work on MD2C platforms makes hit probability depend on the competitive set and the number of solicited dealers, "how many dealers are asked" being a first-order primitive of RFQ economics [5].
The aggregator literature makes the loop explicit: the competitive set is typically a subset of eligible liquidity providers, shaped by ranking, relationship, or "top list + exploration" logic, and "inclusion and rank depend on slowly varying performance scores (e.g., long-run win ratio, response quality), which in turn depend on how aggressively the LP quotes" [24]. Formalising this as a two-tier control problem, aggregator flow whose opportunity intensity is multiplied by a promotion gate driven by the dealer's win score, plus ungated background flow, yields score dynamics that, for steep logistic gates, "can exhibit fold bifurcations, bistability, and hysteresis, producing an endogenous 'campaign vs. harvest' pattern in optimal quoting" [8]. Two consequences follow for training data. First, the support of future data is policy-dependent: quote wide for a while and the promotion gate throttles the aggregator flow that generates labels; quote tight and the observed sample is drawn from a different, more promoted regime [8][24]. Second, hysteresis means the mapping from policy to observed sample is path-dependent, so a naively pooled historical dataset mixes regimes that a single stationary response function cannot represent. Platform design compounds this: the winner's-curse logic gives clients an economic rationale to limit the number of LPs exposed to each RFQ and to use routing and ranking rules rather than inviting the full pool [23], and even absent last look "the aggregator itself shapes adverse selection and incentives" [5].
This is textbook performativity: a framework in which deployed models "recursively influence data distributions through agents' strategic responses to predictions" [15], with performative risk defined over the distribution the model itself induces and performative stability defined as the fixed-point condition that a parameter be optimal with respect to its own induced distribution [14]. Retraining is then not a nuisance but the natural equilibrating dynamic whose fixed points are performatively stable models [17][16], with convergence of repeated risk minimisation guaranteed under strong convexity, smoothness, and sufficiently small Lipschitz sensitivity of the distribution map [12][19][18]; a stable model "achieves minimal risk for the distribution it induces and hence eliminates the need for retraining" [11][13]. Crucially, stability is not optimality, performatively stable points may be suboptimal in performative risk, and the gap is generally non-zero [12].
Two failure modes are directly relevant to a quoting policy that shapes its own labels. First, confounding: because dealers set spreads conditioning on the market environment, RFQ characteristics, instrument, client identity, and the competitive landscape, including the number of dealers in competition and whether the dealer has an axe [30], historical spread/outcome correlations "may not reflect causal relationships" and can bias optimal pricing [7]; identification requires a back-door-admissible conditioning set (volatility plus RFQ, bond, and client features in the cited framework) [7], and where the effect is not identifiable from history, randomised controlled trials / A-B tests are needed [40]. Second, sample-versus-population performativity: the more a model is used to intervene, the more the sample deviates from the original population, "paradoxically making it harder to reliably infer population properties" [35]. Formally, the two directions of failure are a self-negating population (a min-max, distributionally robust functional over a Wasserstein ball) and a self-fulfilling sample, "empirical echo chambers" corresponding to a min-min, distributionally favourable functional, with generalisation bounds driven by covering numbers, Wasserstein sensitivity, and an observable performative response rate
By loose analogy, the generative-AI literature documents "model collapse," a degenerative process in which models trained on their predecessors' outputs progressively forget the true underlying distribution, with tails vanishing first [44], attributed to functional-approximation, sampling, and learning errors [43][46], and characterised as a narrowing and distortion of the output distribution over successive generations [45][42]. The analogy is imperfect, that literature concerns generative models trained on uncurated synthetic data, not dealer quoting, and the corpus contains no study transferring it to RFQ pricing; it is offered only as an intuition for why iterated self-training on one's own induced sample degrades tail coverage.
Finally, evaluation inherits the problem. A quantum-feature study of European corporate-bond RFQ fill-probability estimation notes that any performance gain from a new fill-probability estimator "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs"; backtes
Can realistic market simulators or agent-based/generative models of OTC bond flow support fixed-point training and counterfactual policy evaluation?
The institutional target is well specified by the corpus. Corporate bond execution is dominated by multi-dealer-to-client (MD2C) request-for-quote (RFQ) protocols in which a client solicits a small panel of dealers and trades, if at all, with the best respondent [1][2][4]. Consequently the dealer's win probability is "not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], and hit probabilities depend explicitly on the competitive set and on the number of solicited dealers, making "how many dealers are asked" a first-order primitive of RFQ economics [5]. A simulator that is to be used for fixed-point (self-consistent) training must therefore endogenise at least: rival quoting, client reservation prices, panel composition, and the dealer's own inventory feedback [1][5][6].
Two further structural features constrain fidelity. First, observability is partial: the winner learns the cover price, while losing dealers learn only their ranking and whether a trade occurred, and in some cases cannot even distinguish a missed from a passed RFQ [4]. Second, endogenous dynamics are pervasive: empirical research cited in the quantum-features study suggests many instances of market impact and price jumps arise "without any association to prior events," alongside sparse, statistically irregular, high-dimensional inquiry data and second-scale response windows [32]. Both features limit how tightly a simulator's rival-quote and impact components can be identified from dealer-side data.
The corpus supplies three usable building blocks, with different degrees of endogeneity.
-
Structural/generative RFQ models. Fermanian et al. estimate dealer-quote and client-reservation distributions and derive ex-ante hit-ratio functions [1], and the MD2C causal graphical model specifies a joint distribution over dealer pricing policy, competitors' pricing policies, client reservation price, axes, volatility and revenue definitions [29][30]. A generative specification of this kind models the joint distribution and, once estimated, "can be used to compute any conditional probability through standard probabilistic rules" [38], which is precisely the property a simulator needs. Revenue can be scored under alternative definitions (end-of-day flow value with mid-price inventory valuation and optional liquidity penalties; short-term flow value over seconds-to-minutes windows) [29], although hedging activity is explicitly not included and would require extension [29].
-
Stochastic-control environments. The hit-ratio-targeting model specifies mid-price diffusion, size-ladder quoting, quote-dependent fill probabilities, point-process fills, inventory/cash dynamics, and a size-weighted instantaneous hit ratio with tier-specific targets and penalties [34]; RFQ opportunity arrival intensities there are exogenous, with only the fill probability endogenous to the quote [34]. This is exactly the reduced-form treatment that the same paper criticises: most of the literature "treats execution probability in reduced form, typically through quote-dependent intensities or fill probabilities" [21]. Such an environment therefore embeds an exogenous-flow assumption at the opportunity level even while being endogenous at the fill level.
-
Explicitly performative environments. The two-tier aggregator model closes the loop: aggregator-tier opportunity intensity is multiplied by a promotion gate driven by the dealer's slow win score, and tier-A outcomes update that score, so today's quoting changes tomorrow's flow [8][24]. This is a genuine fixed-point/performative environment, and its behaviour is a warning as much as an enabler: with steep logistic gates the score dynamics exhibit fold bifurcations, bistability and hysteresis, producing an endogenous "campaign vs. harvest" pattern in optimal quoting [8].
-
Agent-based and multi-agent RL models. A parallel strand uses reinforcement learning and agent-based models to study tacit coordination in electronic markets and in MD2C-type dealer platforms, and the role of heterogeneity in mitigating it [23][24], including work on algorithmic collusion in MD2C platforms [22]; other work moves to explicit Stackelberg/Nash equilibrium characterisations of competitive liquidity provision [22][23]. Note that in the corpus these are used to study mechanisms (coordination, equilibrium existence), not as validated training environments with reported policy-transfer results.
So: yes, the corpus supports the claim that fixed-point training environments for OTC bond/RFQ flow are constructible, and that at least one explicitly performative RFQ model exists [8]. It does not contain any demonstration that a policy trained inside such a simulator was deployed and outperformed a baseline in live RFQ flow; that evidence is missing.
The corpus's strongest evidence concerns the identification layer rather than the simulator layer. Because dealers set spreads using market environment, RFQ, instrument, client and competitive information, historical spread–outcome correlations "may not reflect causal relationships" [7][9], and pricing decisions require the interventional distribution p(win | do(spread), ·) rather than the observational conditional [10]. The back-door criterion identifies a minimal conditioning set, volatility plus RFQ, bond and client features, under which standard estimators recover the causal hit-probability function [7], and omitted variables can be handled by integrating over their distribution [7]. Conditioning on all available features is harmful: it can introduce spurious correlations and distort causal-effect estimates [10]. Ignoring these confounders, bond and client characteristics, partially observable factors such as information asymmetry or price-discovery intent, biases predictive models and yields suboptimal decisions [27][30]. Censoring of unwon RFQs is handled by framing the target as won vs. lost/passed, which "enables estimation of trading probability without conditioning on observing actual transactions" [33]; feature selection combines the back-door minimal set with domain knowledge and random-forest importance [33]. Crucially, where the interventional effect is not identifiable from historical data, the corpus's prescription is experimentation, randomised controlled trials implemented as A/B tests, not simulation [40].
The corpus is explicit that direct validation of a policy change is not available: any performance gain from embedding a different fill-probability estimator "can only be assessed indirectly, since trading behaviour changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs" [39]. This is the counterfactual-evaluation problem stated in the dealer's own terms. The fallback is backtesting with a fixed test environment, used for relative comparison of models rather than absolute magnitudes: rolling training windows, in-sample cross-validation, out-of-sample evaluation on the next RFQ with an explicit "blinding window" to prior market information [39]. The same source concedes that, because train/test are strictly decoupled by market time and the data are realisations of stochastic processes, "no general and data-instance independent theoretical guarantees on prediction errors can be made for such out-of-sample or out-of-distribution tests," motivating a purely empirical approach [41].
The corpus contains no stylised-fact battery, no calibration/discriminator-based fidelity test for an OTC bond simulator, and no reported sim-to-real transfer study. Model-selection evidence is limited to the generative-vs-discriminative comparison framing [27][38] and the general statement that discriminative models "often outperform generative models in predictive tasks" and benefit from regularisation and ensembles that "improve out-of-sample generalization" [38], the numerical results of that comparison are not in the corpus.
Fixed-point training loops that iterate on model-generated data inherit the model-collapse literature's failure mode: training on uncurated synthetic data or on the outputs of prior model versions degrades models [42][45], causes them to "forget the true underlying data distribution, even in the absence of a shift in the distribution over time," starting with disappearing tails and converging to point estimates with very small variance [44]. Mechanisms include functional approximation error, sampling error and learning error, compounding in more complex models [43][46], with finite-sampling bias producing increasingly peaked distributions [46]. Access to real data is essential where tails matter [44], a direct concern for RFQ pricing, where the economically decisive events (large sizes, illiquid bonds, informed clients) are tail events [29][32].
Performative prediction formalises the loop in which deploying a model shifts the distribution it is evaluated on [15][17]. A performatively stable model minimises risk on the distribution it induces, so retraining returns the same parameters; stable points are fixed points of risk minimisation [11][14][16], and they "achieve minimal risk for the distribution they induce and hence eliminate the need for retraining" [13]. Retraining is thus reinterpreted as the natural equilibrating dynamic rather than a nuisance [17][16]. Convergence results: repeated risk minimisation converges to a unique stable point when the sensitivity parameter is small enough [18], with linear convergence under strong convexity of the loss, smoothness, and Lipschitz sensitivity of the distribution map [12][19]; weak convexity alone is insufficient [16]; repeated gradient descent achieves the same without an exact optimisation oracle, though over a smaller parameter range [18][16].
Applied to a dealer, the mapping is direct: quoting policy → win/loss outcomes → win score → platform routing → future opportunity intensity [8][24], and reinforcement learning is itself a case of performative prediction because the policy changes the state–action distribution [16].
The corpus supports a conditional and largely theoretical answer, not an empirical one.
- In favour, on identification grounds: policies built on back-door-identified, confounder-controlled response functions avoid the documented bias of associative spread–outcome relations [7][10][27], and discriminative estimators with regularisation/ensembling are described as improving out-of-sample generalisation [38].
- In favour, on equilibrium grounds: a stable policy is by construction optimal for the flow it induces, removing the chase-the-moving-target dynamic [11][13][17].
- Against a strong claim: stability is not optimality. The performatively stable point is merely a fixed point of RRM and "may be suboptimal in terms of performative risk," while the performative optimum need not be stable; the gap is generally nonzero [12][11]. So a fixed-point-trained policy can be systematically dominated by a policy that deliberately exploits its own performative effect.
- Missing: the corpus contains no head-to-head empirical comparison, in bonds or elsewhere, of a performative/fixed-point-trained quoting policy against one trained under an exogenous-DGP assumption. One corpus entry's claim asserts that performative-aware policies "achieve better out-of-sample performance," but its supporting evidence states only that performative drift is incorporated into generalisation bounds via covering numbers and Wasserstein distances [36]; that is a bound-construction result, not a demonstrated performance ranking. This should be read as a conflict between the summarised claim and the underlying evidence, resolved in favour of the weaker reading.
- Strong performativity / weak regularity. RRM converges rapidly only when the stability parameter is below one; when performative effects are strong or regularity weak, fixed-point computation becomes computationally hard, PPAD-complete even in quadratic and linear settings [12].
- Multiple equilibria and path dependence. In the aggregator score model, steep promotion gates generate fold bifurcations, bistability and hysteresis [8]; where multiple attractors exist, "the" fixed point is not well defined and the trained policy depends on initialisation and history, producing campaign-vs-harvest regime switching [8]. Background (non-gated) flow plays a stabilising role by preserving inventory-mixing capacity when the dealer is weakly promoted [8][24].
- Heavy intervention on small samples. Performative learning theory identifies an inherent tension between intervention and inference: "the more a model is used to intervene in the data distribution, the more the sample deviates from the original population, paradoxically making it harder to reliably infer population properties" [35]. Excess-risk bounds degrade with the observable performative response rate m/n [37]. Two distinct failure directions are characterised: population self-negation (worst case = a min–max/DRO functional over a Wasserstein ball) and sample self-fulfilment (retraining ≈ min–min/DFO, i.e. an empirical echo chamber) [37]. The echo-chamber direction is the theoretical counterpart of the model-collapse mechanism above [44][46]: a fixed-point loop can converge to a self-confirming, low-variance distribution rather than to the truth.
- Weak or no performativity. Where flow is genuinely exogenous, reduced-form quote-dependent fill intensities are adequate by construction [21][34], and the extra machinery buys nothing beyond what confounder control already provides [7].
- Non-identifiable interventions. Where do-calculus fails on historical data, no amount of simulator training substitutes for randomised experimentation [40].
The corpus supports the claim that (i) generative causal-graph and stochastic-control models of OTC RFQ flow are rich enough to serve as fixed-point training environments [29][30][34][38], (ii) at least one explicitly performative RFQ environment with score→flow feedback exists and is solvable [8], (iii) counterfactual pricing/revenue questions on MD2C platforms have been studied via causal interventions bridging structural RFQ models and discriminative learning [5][27], and (iv) fidelity is currently validated only weakly, by time-decoupled backtesting with explicit acknowledgement that policy-induced market responses are unobserved [39][41]. What the corpus does not provide: quantitative fidelity metrics or acceptance tests for OTC bond simulators; calibration of rival-quote distributions under the corpus's own partial-observability constraints [4]; off-policy evaluation estimators; live A/B results for RFQ pricing despite the framework recommending RCTs [40]; and any measured out-of-sample performance advantage of performative-aware or causally-invariant policies over exogenous-DGP-trained baselines. On present evidence, the case for fixed-point/causal training rests on identification arguments [7][10][27] and equilibrium logic [11][13][17], tempered by the stability–optimality gap [12], hardness and multiplicity under strong feedback [12][8], and echo-chamber/model-collapse degeneration in self-consuming training loops [37][44][46].
Any credible test protocol has to begin by pinning down the object that is supposed to be a fixed point. The performative-prediction literature gives two distinct definitions, and they demand different experiments. A model is performatively stable if it minimises expected loss on the distribution that its own deployment induces, so that any retraining procedure would simply return the same parameters, stable models are fixed points of risk minimisation [11], formally the fixed-point condition that a parameter is optimal with respect to its own induced distribution [14], and equivalently the fixed points of the retraining (RRM) operator [12], [16]. A model is performatively optimal if it globally minimises performative risk, which need not be a fixed point [12]. The two concepts are in general distinct, and stable points can be suboptimal in performative risk, a gap described as central both to the theory and to the practical pathology of performative settings [12].
In the corporate-bond RFQ setting, the candidate fixed point is a quoting policy whose implied hit probabilities are consistent with the flow it actually attracts. The corpus supports treating this as a genuinely endogenous object: the dealer's hit ratio "is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], hit probabilities depend on the competitive set and on the number of solicited dealers [5], and in industry practice historical hit ratios are themselves used to select dealers and route RFQs [21]. That last fact is the closed loop: quotes → wins → measured hit ratio → future RFQ invitations → future data.
Two testable implications follow. First, retraining invariance: at a fixed point, refitting on freshly induced data should reproduce the deployed policy [11], [13], [16], [17]. Second, convergence behaviour: RRM-type iteration converges linearly to a unique stable point under strong convexity, smoothness, and Lipschitz sensitivity of the distribution map, with contraction governed by a stability parameter combining sensitivity, smoothness and strong convexity [12], [13], [19]; weak convexity alone is insufficient [16], and gradient-based approximations of retraining converge over a strictly narrower range of sensitivity than exact RRM [16], [18]. When sensitivity is large, fixed-point computation becomes computationally hard (PPAD-complete), even in quadratic and linear settings [12]. So the sensitivity/response magnitude is not a nuisance parameter, it is a primary estimand.
Observational quote data cannot identify the response function, because dealers do not set spreads from context-independent drivers: they incorporate the market environment, the RFQ itself, instrument characteristics, client identity and the competitive landscape, which "introduces correlations between historical spreads and RfQ outcomes that may not reflect causal relationships" [7], [3]. The causal graph makes the confounding explicit: client identity, bond features (liquidity, relative value), RFQ features (number of dealers in competition) and axes all feed the pricing policy and the outcome, creating spurious dependence between quoted price and RFQ result [30], and volatility influences both dealer pricing policies and the client's reservation price [29]. Optimal pricing therefore requires the interventional distribution under do(spread), not the historical conditional [10], and the back-door criterion identifies the minimal conditioning set, volatility plus RFQ, bond and client features [7]. Where the intervention effect is not identifiable from history, the corpus is explicit that one must "design experiments, specifically, randomized controlled trials (RCTs), commonly implemented as A/B tests, to empirically measure the impact of the intervention" [40].
Design implications:
- Randomise the offset, stratify on the back-door set. Randomising a perturbation to the policy spread severs the confounding arrows into the price node [30], [10]; stratifying by volatility, RFQ, bond and client features preserves the conditioning set needed for valid estimation and for later re-use of the estimates [7].
- Respect the natural cell structure. The control problem is indexed by bond, client tier, side, and size-ladder rung, each with its own arrival intensity, with size-weighted instantaneous hit ratio defined per tier and explicit target levels per targeted tier [34]. Randomisation and estimation should be organised on the same grid, and the model itself notes that differentiated targets across instrument subsets are implemented by splitting tiers [34], the same device gives clean experimental strata.
- Pre-specify the revenue metric. The corpus documents at least three inequivalent revenue definitions: realised spread, end-of-day flow value using mid-prices to value inventory (optionally with liquidity penalties on unrealised mark-to-market) and short-term flow value over a seconds-to-minutes window, the last being useful for detecting clients trading on asymmetric information before external factors dominate [29]. Since dealers must also trade off win probability against margin, adverse selection and post-trade market movements [3], [4], a single headline hit-ratio metric is insufficient; hit ratio, margin, and mark-out at short and end-of-day horizons must be reported jointly.
- Handle censoring honestly. Feedback is only partially observable: the winner learns the cover price, others learn ranking and whether a trade occurred, and in some cases cannot distinguish a missed from a passed RFQ [4]. This argues for the target used in the empirical work in the corpus, a binary won/lost classification that "enables estimation of trading probability without conditioning on observing actual transactions" [33], and for treating unobserved outcome status as a documented data limitation rather than silently dropping it. Failure to account for partially observable factors such as information asymmetry or price-discovery intent is shown to bias predictive models and produce suboptimal decisions [27].
- Governance and auditability. Experiments live inside a regulated workflow: MiFID II requires firms to take sufficient steps to achieve best execution and to maintain evidence, supported by pre-trade price comparison across counterparties, a market-conditions snapshot at execution, and post-trade TCA [48], [47], with detailed audit trails an explicit driver of electronification [49]. Venue-level observability is also increasing on the sell side, e.g. FINRA's requirement to report the ATS MPID for TRACE-eligible securities [51]. Model risk management and regulatory controls are cited as practical constraints on model complexity in this exact market [50]. The corpus does not address whether deliberately randomising client-facing prices is permissible under these regimes; that evidence is missing and should not be assumed either way.
Randomised experiments recover the response function at a point; they do not by themselves test the fixed-point claim. The natural test mirrors the theory: deploy θ_t, collect the induced flow, refit, and track the parameter movement across deployment rounds, since fixed points of retraining are the performatively stable points [16] and a stable model eliminates the need for retraining because retraining returns the same parameters [11], [13]. Falsification criteria:
- Non-convergence or oscillation of the deployment sequence is evidence that effective sensitivity is outside the contraction regime [12], [18], recalling that the gradient-style approximation converges on a narrower range than exact risk minimisation [16].
- Convergence rate is itself an estimate of the stability parameter, since RRM converges linearly with a rate set by sensitivity, smoothness and strong convexity [12].
- Stability ≠ optimality must be tested separately. Because a fixed point may be suboptimal in performative risk [12], the protocol needs a long-horizon comparison between the stable policy and candidate performative-optimum policies, evaluated on the distribution each one induces [12], [13], [19]. Note the theoretical caveat that existence of unique stable points was established under strong convexity and smoothness, with weaker-loss existence results only in constrained solution spaces [19], and existence of stable points at all requires a sufficiently Lipschitz distribution map [13].
- Measure the response rate. Recent performative learning theory identifies a directly observable quantity: the performative response rate m/n, the proportion of units in the sample that changed because of the prediction, with generalisation bounds growing in m [37]. Operationally, this maps to the share of RFQs whose invitation, competition, or outcome changed as a consequence of the deployed policy, the single best summary statistic for "how performative is this book".
The corpus contains an explicit backtesting protocol for exactly this market: at each RFQ arrival, build a training window of past events, cross-validate model variants, select the best in-sample model, and test out-of-sample on the next RFQ, with a deliberate blinding window between the last training information and the evaluation instance, iterating forward [39]. It also insists on strict temporal decoupling of train and test by market time, notes that this makes the inputs random variables subject to the underlying stochastic processes, and concludes that "no general and data-instance independent theoretical guarantees on prediction errors can be made for such out-of-sample or out-of-distribution tests", motivating a purely empirical design [41].
Crucially, the same source states the limitation that matters here: any performance gain from embedding a fill-probability estimator "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs", backtesting is defended only as a relative comparison "while keeping the test environment fixed", and the paper is explicit that it studies observable differences under fixed model parameters, not the accuracy of their magnitude [39]. That is a direct admission that a fixed-environment backtest cannot test a fixed-point hypothesis: the hypothesis is about how the environment moves. Backtests can falsify calibration and discrimination; they cannot validate equilibrium claims.
Two partial substitutes appear in the corpus. First, counterfactual causal evaluation on MD2C data, causal interventions have been used to study counterfactual pricing and revenue questions, bridging structural RFQ models with discriminative learning [5], [27]. Second, simulation grounded in structural quoting models that combine inventory-sensitive pricing with a realistic description of how RFQs translate into fills [6], including quote-dependent fill probabilities, size ladders and explicit hit-ratio penalties [34], [21]. Both remain model-dependent; neither is presented in the corpus as validated against a live randomised benchmark.
Also note a methodological trap specific to backtests here: feature selection driven purely by predictive importance risks spurious
Adaptive quoting in fixed income does not operate in a rule-free environment. The most binding documented constraint is the best-execution regime: under MiFID II, buy-side firms must "take sufficient steps to achieve best execution and maintain evidence that they have done so" [48], and this obligation has been a direct driver of electronification because "manual, voice-based workflows cannot easily satisfy at scale" the requirement to "demonstrate best execution and maintain detailed audit trails" [49]. Operationally, compliance is discharged through the execution management stack: automated pre-trade price comparison across multiple counterparties, capture of "a complete snapshot of market conditions at the time of execution," and post-trade transaction cost analysis, replacing spreadsheet processes with "automated, auditable records" [48]. Pre-trade transparency is a core EMS capability precisely so that "traders need to know whether they are seeing a fair price before they trade," with side-by-side price comparisons benchmarked against live market data [47]; TCA then lets clients "measure performance against benchmarks, understand dealer relationships, and identify opportunities to improve," data that is "critical for demonstrating best execution under regulatory frameworks such as MiFID II" [47].
On the market-integrity side, FINRA has flagged fixed-income-specific concerns: "increased trading of TRACE-eligible securities in dark pools, greater use of request-for-quote processes, and more executions of customer orders by firms that participate as both broker and as dealer," prompting a requirement that firms report the unique MPID of the ATS to TRACE [51]. FINRA also states explicitly that it is "concerned about some of the protocols within the platforms" [51], that is, the RFQ mechanism itself, not merely the prices it produces, is a supervisory object. A third, less visible constraint is internal: model risk management and regulatory controls create pressure on model complexity, which is why at least one industry-oriented approach deliberately confines innovation to the input features so that "no changes to the currently used learning algorithm would be a priori required, and a potentially less complex or more robust model may be chosen, which keeps the model complexity manageable" [50].
A fixed-point (performatively stable) quoting policy must be optimal for the distribution it itself induces, and the canonical way to find one is repeated retraining on data generated by the previous deployment [14][16]. That procedure presupposes the ability to deploy and observe consequences. The causal-inference literature on dealer pricing makes the requirement sharper: because dealers "actively incorporate information from the market environment, the RfQ itself, the characteristics of the instrument, the identity of the client, and the competitive landscape … when setting prices," historical spread–outcome correlations "may not reflect causal relationships" [7], and these variables act as confounders that "introduce spurious dependencies between the quoted price and the RfQ outcome" [30]. Standard probability theory is described as "not well suited to address interventional questions," and where do-calculus cannot identify an effect from historical data, "it becomes necessary to design experiments, specifically, randomized controlled trials (RCTs), commonly implemented as A/B tests, to empirically measure the impact of the intervention" [40]. Deliberate quote randomization is therefore not a modelling luxury; it is the fallback when identification fails.
Each element of the compliance perimeter narrows that fallback:
- Auditability makes experiments legible and permanent. Every quote sits inside a client-side record of "a complete snapshot of market conditions at the time of execution" plus post-trade TCA against benchmarks [48][47]. An exploratory quote that is deliberately off-market is, by construction, captured in the client's best-execution evidence and dealer-performance analytics [47][48].
- Counterparty selection converts exploration into lost flow. TCA is explicitly used to "understand dealer relationships" [47], and in industry practice "historical hit ratios are used to select dealers and route RFQs, especially for liquid electronic bond orders, alongside axes, response quality, and relationship information" [21]. Aggregator-routed markets go further: "the competitive set is typically a subset of all eligible LPs and is often shaped by ranking, relationship, or platform-level 'top list + exploration' logic," with inclusion and rank depending on "slowly varying performance scores (e.g., long-run win ratio, response quality), which in turn depend on how aggressively the LP quotes" [24]. Exploration that degrades the score therefore reduces the learner's future data supply, the platform, not the dealer, controls the exploration budget.
- Platform design deliberately limits who may compete. Oomen's winner's-curse argument gives "an economic rationale for limiting the number of LPs exposed to each RFQ and for using routing and ranking rules rather than always inviting the full pool" [23]; the number of dealers asked is a "first-order primitive in RFQ economics" [5]. A dealer cannot self-serve additional RFQ observations.
- Feedback is censored. In MD2C protocols "the winning dealer is informed of the second-best quote (the 'cover price'), while the other dealers receive feedback on their ranking and whether a trade occurred," and in some cases losing dealers "cannot even know if the trade occurred (a missed RfQ) or not (a passed RfQ)" [4]. Exploratory quotes thus return weak or partially missing labels, raising the number of experiments needed for a given inferential gain.
- Exploration has direct P&L and risk cost. Quoting aggressively "may increase hit probability but reduce margins or expose the dealer to adverse selection and post-trade market movements" [3][4], while inventory risk and hit-ratio deviation penalties are formalized as explicit objective terms in dealer control problems [34].
- Exploration can be irreversible. Where routing depends on a slow win score with a steep promotion gate, "the score dynamics can exhibit fold bifurcations, bistability, and hysteresis, producing an endogenous 'campaign vs. harvest' pattern in optimal quoting" [8]. Hysteresis means an exploratory episode is not a costless perturbation that decays: it can move the dealer between regimes.
Even a successful fixed point raises integrity questions, because rivals are also learners. A parallel strand of research "uses reinforcement learning and agent-based models to study the emergence of tacit coordination in electronic markets and in MD2C-type dealer platforms, and the role of heterogeneity in mitigating such effects" [24][23], including work on "algorithms and supracompetitive prices in electronic markets" and "algorithmic collusion in multi-dealer-to-client platforms" [22]. Since the dealer's hit ratio "is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], a policy that is stable against the induced distribution may be stable because rival algorithms have co-adapted, exactly the pattern FINRA's concern with "protocols within the platforms" would target [51]. Strategic-interaction models (Stackelberg, Nash) are the formal apparatus for characterizing such equilibria [23][22], but the corpus contains no regulatory guidance specifying when a learned equilibrium crosses from competition into coordination; that mapping is missing from the evidence.
Client identity is observable on MD2C platforms and is used in pricing: "since the client identity is known by the dealers quoting in MD2C platforms, their pricing policies tend to incorporate specific features of clients" [30], and the causal framework notes that price segmentation, "when feasible, … can yield higher revenues than global pricing strategies, as extensively discussed in classical microeconomics and marketing literature" [7]. The same source is careful that accurate hit-probability estimation "does not imply that the dealer must adopt price segmentation strategies" [7]. This is the sharpest fairness tension in the corpus: the modelling literature supplies an efficiency argument for differential, client-conditional quoting [7][30], while the compliance material speaks only to the client's obligation to obtain best execution and evidence it [48][49], it says nothing about limits on dealer-side price discrimination or about disclosing to clients that their RFQs may be used as randomized treatment arms. That evidence is absent from the corpus and should not be assumed either way. Separately, the performative-prediction literature lists fairness alongside robustness and computational complexity as open challenges in applications such as dynamic pricing [15], but the corpus provides no fairness formalism specific to RFQ quoting.
Because live experimentation is constrained, practitioners lean on three substitutes, each with limits evidenced here.
- Identification instead of randomization. The back-door criterion "provide[s] a principled way to identify the minimal set of variables required for valid causal effect estimation" [7], with volatility, RFQ, bond and client features forming the minimal conditioning set, and with integration over omitted variables when the conditioning set is incomplete [7]. This reduces reliance on RCTs but only under a correctly specified causal graph; the framework's own conclusion warns that "neglecting confounders … or failing to account for partially observable factors, including information asymmetry or client intent (e.g., price discovery), can bias predictive models and ultimately result in suboptimal decision-making" [27].
- Backtesting. Backtesting is "widely used in model validation in the financial industry" and "lets us relatively compare and benchmark different models while keeping the test environment fixed" [39]. Its stated limitation is precisely the performative channel: any gain "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market that influence the conditions for the next time instance and respective RFQs" [39]. Moreover, with strict temporal separation of train and test sets, "no general and data-instance independent theoretical guarantees on prediction errors can be made for such out-of-sample or out-of-distribution tests" [41]. Backtests therefore cannot certify a fixed point; they can only rank candidates in a counterfactually frozen world [39][41].
- Simulation and synthetic data. The corpus documents a generic hazard for this route: "indiscriminately learning from data produced by other models causes 'model collapse', a degenerative process whereby, over time, models forget the true underlying data distribution" [44], driven by functional-approximation, sampling and learning errors [43] and by finite-sampling bias producing increasingly peaked distributions [46], with outputs becoming "increasingly generic, repetitive, and disconnected from the true distribution" [45]. Applying this to a self-simulated RFQ environment is an extrapolation the corpus does not test directly, but the mechanism is documented [43][44][45][46].
Even absent any rulebook, restraint on exploration has a statistical rationale. Performative learning theory frames the core difficulty as "intervention vs. inference": "the more a model is used to intervene in the data … the more the sample deviates from the original population, paradoxically making it harder to reliably infer population properties" [35], with generalization bounds that grow in the "performative response rate m/n", the share of sampled units that changed because of the prediction [37]. The theory further identifies two failure directions: population self-negation, whose worst case corresponds to distributionally robust optimization over a Wasserstein ball, and sample self-fulfilment, where empirical retraining is equivalent to distributionally favourable optimization and produces "empirical echo chambers" [37]. Practically, this means the compliance-driven cap on exploration is not purely a cost: bounds on out-of-sample performance degrade with the intensity of intervention [35][37], and performative-aware objectives that price drift explicitly are the proposed remedy [36]. It also means the stable point the learner reaches may be the wrong target: performatively stable points are fixed points of retraining that "may be suboptimal in terms of performative risk," and the stability–optimality gap "is generally nonzero" [12][11]. Convergence itself requires strong convexity, smoothness and sufficiently weak performative sensitivity, and when sensitivity is large "fixed-point computation becomes computationally hard (PPAD-complete)" [12][16][19].
The corpus supports the existence and mechanics of best-execution and audit-trail obligations [48][49][47], ATS/TRACE reporting and protocol-level supervisory interest [51], and model-risk/complexity pressure [50]. It does not contain: (i) any rule text or supervisory guidance governing algorithmic testing, kill switches, or pre-deployment validation of adaptive quoting engines; (ii) any regime specifying permissible randomization of client-facing prices, or client-consent/disclosure requirements for pricing experiments; (iii) any fair-dealing, markup, or anti-discrimination standard applicable to counterparty-conditional bond quoting; (iv) any quantitative regulatory or platform-imposed exploration budget; and (v) any documented case of a regulator assessing an algorithmic-collusion outcome on an MD2C platform. Claims in those areas cannot be substantiated from this evidence base.
The corpus supports a coherent chain of reasoning. First, the institutional setting: corporate bond trading remains predominantly over-the-counter and dealer-intermediated, with electronification proceeding through multi-dealer-to-client request-for-quote (RFQ) protocols rather than centralized limit-order books, and clients typically trading with the best respondent among a small solicited panel [1]. RFQ is described as the most widely used electronic execution protocol in bond markets [2], and MD2C platforms as the dominant architecture for institutional bond trading [4]. Within this structure, the dealer's problem is jointly one of inventory management, counterparty selection and endogenous execution probability: quoting must be competitive enough to win flow without destroying margin or accumulating unwanted inventory [1], [3], [4].
Second, the endogeneity is explicit in the literature rather than merely implied. Execution is shaped not only by price but by the protocol itself, the number of requested dealers, competitor response behaviour, and client search technology [1]; the number of solicited dealers is characterised as a "first-order primitive in RFQ economics" [5]. The synthesis offered is that "the dealer's hit ratio is not an exogenous statistic, but an equilibrium object shaped by platform design, client search, and rival participation" [1], [6]. On aggregator-routed venues, the feedback runs through platform machinery: routing intensity is gated by a slowly varying win score that is itself a function of how aggressively the dealer quotes, and with steep promotion gates the score dynamics can exhibit fold bifurcations, bistability and hysteresis, producing an endogenous "campaign vs. harvest" quoting pattern [8]; even absent last look, the aggregator shapes adverse selection and incentives [5].
Third, this is formally the structure of performative prediction: a framework in which deployed models recursively influence the data distribution through agents' strategic responses [15], with performative stability defined as the fixed-point condition that a parameter be optimal with respect to its own induced distribution [14], [11]. Retraining is reinterpreted not as a nuisance response to distribution shift but as the natural equilibrating dynamic whose fixed points are performatively stable models [17], [16]. Two conclusions follow directly for RFQ pricing. (i) A hit-probability model estimated on data generated under a prior quoting policy is a moving target, because the dealer's own quotes enter the data-generating process, a point the causal literature makes independently by insisting on the do-operator: what matters is the interventional distribution of wins given spreads, not the historical conditional distribution [10]. (ii) The mathematics of convergence is conditional, not automatic: RRM and repeated gradient descent converge linearly to a unique stable point under strong convexity, smoothness, and sufficiently small Lipschitz sensitivity of the distribution map [12], [13], [18], [19], with weak convexity alone insufficient [16]; when performative effects are strong or regularity weak, fixed-point computation becomes computationally hard (PPAD-complete) [12].
A recurring finding is that the performative/market-response function cannot be read off naively from dealer logs. Dealers do not quote on context-independent drivers; they incorporate market environment, RFQ features, instrument characteristics, client identity and the competitive landscape, which "introduces correlations between historical spreads and RfQ outcomes that may not reflect causal relationships" [7], [9-adjacent evidence in [30]]. The proposed remedy is a causal graphical model of the RFQ process, with the back-door criterion identifying a minimal conditioning set, volatility plus RFQ, bond and client features, for valid causal estimation of hit probability [7], [3], [27], and with client/bond/competition variables explicitly labelled confounders that would otherwise bias optimal pricing [30]. Where effects are not identifiable from history, the same framework prescribes randomized controlled trials / A/B tests [40]. Feature selection in the reported empirical work is anchored on this minimal conditioning set combined with domain knowledge and random-forest importance [33]. There is a stated methodological trade-off: discriminative models often outperform generative ones predictively and benefit from regularization and ensembling for out-of-sample generalization, but are more opaque and worse at representing structural mechanisms [38].
The corpus is unusually candid that standard evaluation practice is not performativity-aware. In a bond RFQ fill-probability study, the authors state that performance gains from embedding a new estimator "can only be assessed indirectly, since trading behavior changes as a consequence of different fill probability estimates may lead to unknown responses from the market," and that backtesting is used precisely because it holds the test environment fixed [39]; they further argue that no data-instance-independent theoretical guarantees on prediction error are available for such out-of-sample/out-of-distribution tests, motivating a purely empirical protocol [41]. Learning theory work addresses the same gap from the other direction: once predictions react back on the distribution, classical train/test generalization conclusions no longer apply, and the more a model is used to intervene, the more the sample deviates from the population, an inherent "intervention vs. inference" tension [35]. Bounds are obtainable without assuming a functional form for the transition map, using Wasserstein sensitivity, covering numbers and a performative response rate [36], with failure modes characterised as population self-negation (an inf-sup / distributionally robust functional) and sample self-fulfilment (an inf-inf / distributionally favorable functional, corresponding to repeated empirical risk minimization) [37].
- Endogenous hit ratio vs. reduced-form fills. The empirical/structural strand treats the hit ratio as an equilibrium object [1], [5], whereas most stochastic-control market-making literature "treats execution probability in reduced form, typically through quote-dependent intensities or fill probabilities, rather than elevating hit ratio itself to an explicit control target" [21]. One line of work responds by making hit-ratio targeting an explicit penalty in the dealer's objective [34], while others embed hedging and market impact [21], [31] or inventory-sensitive pricing with realistic RFQ-to-fill descriptions [6].
- Stability as sufficient vs. insufficient. Stable models "eliminate the need for retraining after deployment" and achieve minimal risk for the distribution they induce [11], [13]; but stable points "may be suboptimal in terms of performative risk," and the stability–optimality gap is generally nonzero [12], [11]. So a fixed-point-trained quoting policy is self-consistent, not necessarily revenue-maximal.
- Single-agent performativity vs. genuine strategic equilibrium. In strategic classification, the Stackelberg equilibrium coincides exactly with the performative optimum [25], suggesting single-agent performativity suffices when rivals are passive. But dealer markets feature explicitly strategic rivals, modelled via Stackelberg/Nash equilibrium characterisations of competitive liquidity provision and via RL/agent-based studies of tacit coordination and algorithmic collusion on MD2C platforms [22], [23], [24]. Reinforcement learning is itself framed as a case of performative prediction [16]. The corpus does not resolve when a single-dealer performative fixed point coincides with a multi-dealer equilibrium, this remains an open question.
- Feedback-loop degradation. The strongest direct evidence on self-referential training loops concerns generative AI: training on model-generated or uncurated synthetic data causes "model collapse," a degenerative loss of the true distribution driven by functional approximation, sampling and learning errors, with tails disappearing first and outputs converging to low-variance point estimates [42], [43], [44], [45], [46]. This is a suggestive analogy for policies retrained on their own censored win/loss data; the corpus contains no direct evidence of model collapse in RFQ pricing systems, and the analogy should not be presented as established.
Deployment sits inside a compliance envelope: MiFID II requires firms to take sufficient steps to achieve best execution and retain evidence, supported in practice by pre-trade multi-counterparty price comparison, market-condition snapshots and post-trade TCA [48], [47], [49]; regulators have separately flagged increased RFQ usage and ATS opacity in TRACE-eligible fixed income, requiring ATS identification in trade reports [51]. Model risk management and regulatory controls are cited as reasons to prefer interventions that change inputs rather than the learning algorithm, keeping model complexity manageable [50].
Notable gaps in the corpus: (a) no empirical estimate of the magnitude of the Lipschitz sensitivity parameter (and hence of whether the contraction condition of [12], [13] actually holds) in bond RFQ data; (b) no reported production deployment of fixed-point/performatively-stable training in a dealer quoting engine, nor an A/B test of it against a conventional retrained model, the corpus establishes the theoretical desirability [11], [16] and the causal identification machinery [7], [27] but not the empirical payoff; (c) no evidence on how censoring of lost RFQs (where dealers may not even learn whether a trade occurred [4]) interacts quantitatively with performative bias, though the informational asymmetry itself is documented [4], [32]; (d) no evidence that regulators currently treat performative feedback as a distinct best-execution or model-risk concern.
The evidence supports treating dealer RFQ quoting as a performative prediction problem: the quote is an intervention whose distributional consequences run through client behaviour, rival responses and platform routing scores [1], [5], [8], [10]. Consequently, hit-probability models should be identified causally rather than associatively [7], [27], trained with awareness that fixed points exist and are reachable only under sensitivity and convexity conditions that are not guaranteed [12], [13], [16], and evaluated with methods that acknowledge the fixed-environment limitation of standard backtests [39], [41], [35]. The main unresolved issues are the relation between single-dealer performative optima and multi-dealer equilibria [23], [25], the empirical size of performative effects, and whether self-referential retraining in this domain exhibits anything analogous to documented model collapse in generative settings [44], [46].
[1] Bond Market Making with a Hit-Ratio Target - https://arxiv.org/html/2604.20406 (1 Introduction; chunk 89f1a469418e79f8, chars 975-3012) [2] Fixed Income EMS: Electronic Trading in Bonds - https://tsimagine.com/insights/fixed-income-ems-electronic-trading-bond-markets/ (Frequently Asked Questions (FAQs) > What is the difference between RFQ and portfolio trading in a fixed income EMS?; chunk 420428fd5d97a653, chars 15073-15678) [3] Abstract - https://arxiv.org/html/2506.18147v2 (1 Introduction; chunk e8c9d7144c963e2b, chars 3375-5396) [4] Abstract - https://arxiv.org/html/2506.18147v2 (1 Introduction; chunk 91454b1465107000, chars 1711-3649) [5] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (1 Introduction; chunk d606e28c50d53cb4, chars 1816-3678) [6] Bond Market Making with a Hit-Ratio Target - https://arxiv.org/html/2604.20406 (1 Introduction; chunk e9b3d4905e46a126, chars 2595-4495) [7] Abstract - https://arxiv.org/html/2506.18147v2 (3 Causal interventions and predictions in the graphical model > Optimal pricing; chunk dfdf9f5a7080f29c, chars 26194-28096) [8] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (Abstract; chunk 2cdc4ade90a558d8, chars 0-1814) [9] Abstract - https://arxiv.org/html/2506.18147v2 (References; chunk 9a059ed2d16e4724, chars 89241-91279) [10] Abstract - https://arxiv.org/html/2506.18147v2 (3 Causal interventions and predictions in the graphical model > Optimal pricing; chunk de2bcc5ab6395295, chars 24492-26465) [11] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (2 Framework and main definitions > 2.2 Performative stability > Definition 2.3 (performative stability and decoupled risk).; chunk d2e7555ac1613d8d, chars 14554-15426) [12] Performative Prediction: Frameworks & Challenges - https://www.emergentmind.com/topics/performative-prediction (2. Stability, Optimality, and Dynamics; chunk 315b3eb7c8a418ca, chars 2315-3497) [13] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (3 When retraining converges to stable points; chunk 25fdfb5e8753df9d, chars 16373-18242) [14] Performative Prediction: Frameworks & Challenges - https://www.emergentmind.com/topics/performative-prediction (1. Formal Framework and Definitions; chunk 7585101a6b72f832, chars 1202-2313) [15] Performative Prediction: Frameworks & Challenges - https://www.emergentmind.com/topics/performative-prediction (chunk b86b3cdef33cdff9, chars 0-1200) [16] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (6 Discussion and Future Work; chunk dc6baca4f2ecb72c, chars 36307-38249) [17] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (1 Introduction; chunk 332df9246a392059, chars 1145-3175) [18] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (3 When retraining converges to stable points > 3.3 Repeated gradient descent; chunk 050437864d24e4f5, chars 24072-24426) [19] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (4 Relating performative optimality and stability; chunk 9e0596d9bc157a22, chars 27201-27858) [20] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (References; chunk 0cc09d4f2d3423ee, chars 25595-27631) [21] Bond Market Making with a Hit-Ratio Target - https://arxiv.org/html/2604.20406 (1 Introduction; chunk 72fb928960f4d24b, chars 4141-6178) [22] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (References; chunk 3b68bb92ee7f498b, chars 27308-29125) [23] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (1 Introduction; chunk c17a5f66b8f15902, chars 3385-5364) [24] Win-score promotion gates in aggregator-routed RFQ markets: A two-tier stochastic control model - https://arxiv.org/html/2603.10569 (1 Introduction; chunk eaef106a159931cd, chars 4940-6814) [25] Performative Prediction - https://ar5iv.labs.arxiv.org/html/2002.06673 (5 A case study in strategic classification > 5.1 Stackelberg equilibria are performative optima; chunk 0c0bcc8b34cf81f9, chars 30394-32431) [26] Abstract - https://arxiv.org/html/2506.18147v2 (1 Introduction; chunk 0607e5a25e8b3319, chars 4873-6703) [27] Abstract - https://arxiv.org/html/2506.18147v2 (7 Conclusions; chunk a17d436f82075c54, chars 79290-81128) [28] Enhanced fill probability estimates in institutional algorithmic bond trading using statistical learning algorithms with quantum computers - https://arxiv.org/html/2509.17715v1 (1 Introduction; chunk 763c5ec0fd0ed9fc, chars 6173-7920) [29] Abstract - https://arxiv.org/html/2506.18147v2 (2 A causal graphical model for the RfQ process; chunk 6bff77f8cd7aa5d0, chars 18575-20450) [30] Abstract - https://arxiv.org/html/2506.18147v2 (2 A causal graphical model for the RfQ process; chunk 6cbfdc4ef0225470, chars 20159-21381) [31] Bond Market Making with a Hit-Ratio Target - https://arxiv.org/html/2604.20406 (References; chunk 69082f43a159514e, chars 21461-23435) [32] Enhanced fill probability estimates in institutional algorithmic bond trading using statistical learning algorithms with quantum computers - https://arxiv.org/html/2509.17715v1 (2 Methodology > 2.1 Background and motivation; chunk be5f1fa6117c602e, chars 12691-14660) [33] Abstract - https://arxiv.org/html/2506.18147v2 (6 Generative versus discriminative models for causal interventions: empirical results; chunk 5cb81b1aaa443fb2, chars 66227-67973) [34] Bond Market Making with a Hit-Ratio Target - https://arxiv.org/html/2604.20406 (2 OTC market making and hit-ratio targeting > 2.1 State dynamics and objective; chunk 2fe6e6fae94d0724, chars 6721-8615) [35] [Paper Note] Performative Learning Theory - https://en.papernotes.org/ICML2026/learning_theory/performative_learning_theory/ (Background & Motivation¶; chunk f59d8861b63ebf4b, chars 925-2927) [36] [Paper Note] Performative Learning Theory - https://en.papernotes.org/ICML2026/learning_theory/performative_learning_theory/ (Background & Motivation¶; chunk b9780014bbb7622e, chars 2666-3344) [37] [Paper Note] Performative Learning Theory - https://en.papernotes.org/ICML2026/learning_theory/performative_learning_theory/ (Method¶ > Key Designs¶; chunk 32320eea2adf5646, chars 6045-7763) [38] Abstract - https://arxiv.org/html/2506.18147v2 (5 Model specification; chunk 94d1880a1222e7ef, chars 52445-54449) [39] Enhanced fill probability estimates in institutional algorithmic bond trading using statistical learning algorithms with quantum computers - https://arxiv.org/html/2509.17715v1 (3 Empirical analysis setup > 3.3 Trade execution backtesting and benchmarking; chunk bac2d5e8fb7ff156, chars 43672-45607) [40] Abstract - https://arxiv.org/html/2506.18147v2 (3 Causal interventions and predictions in the graphical model; chunk 95a4d62425146e16, chars 21383-22737) [41] Enhanced fill probability estimates in institutional algorithmic bond trading using statistical learning algorithms with quantum computers - https://arxiv.org/html/2509.17715v1 (2 Methodology > 2.4 Trade execution learning; chunk 5c1cfcfc0d1e69d3, chars 25102-26403) [42] Model collapse - Wikipedia - https://en.wikipedia.org/wiki/Model_collapse (chunk 3a1afed95bb0b493, chars 0-491) [43] Model collapse - Wikipedia - https://en.wikipedia.org/wiki/Model_collapse (Mechanism; chunk 6dfa4412a26452d7, chars 1109-1524) [44] AI models collapse when trained on recursively generated data - Nature - https://www.nature.com/articles/s41586-024-07566-y (Main; chunk 9a5ba85990db3f51, chars 1315-3319) [45] Model Collapse: What Happens When AI Trains on AI-Generated Data (2026) - https://aisecurityandsafety.org/en/guides/model-collapse/ (What Is Model Collapse?; chunk e1eea83e4546aaf2, chars 0-1539) [46] Model Collapse Demystified: The Case of Regression - https://arxiv.org/html/2402.07712 (2 Review of Literature > Model Collapse.; chunk c8b629debeb42b82, chars 7337-9218) [47] Fixed Income EMS: Electronic Trading in Bonds - https://tsimagine.com/insights/fixed-income-ems-electronic-trading-bond-markets/ (Core Capabilities of a Fixed Income EMS; chunk 946f9639539d4e2e, chars 2920-4373) [48] Fixed Income EMS: Electronic Trading in Bonds - https://tsimagine.com/insights/fixed-income-ems-electronic-trading-bond-markets/ (Frequently Asked Questions (FAQs) > How does a fixed income EMS help with MiFID II best execution requirements?; chunk 9b04fd94fa387c74, chars 14611-15071) [49] Fixed Income EMS: Electronic Trading in Bonds - https://tsimagine.com/insights/fixed-income-ems-electronic-trading-bond-markets/ (The Electronification of Bond Markets: Why Now?; chunk 7c3732f0542e0c54, chars 4375-6348) [50] Enhanced fill probability estimates in institutional algorithmic bond trading using statistical learning algorithms with quantum computers - https://arxiv.org/html/2509.17715v1 (2 Methodology > 2.1 Background and motivation; chunk 2ffd029378fd6fd1, chars 14256-15980) [51] Remarks at the North American Electronic Bond Trading Forum - https://www.finra.org/media-center/speeches-testimony/remarks-north-american-electronic-bond-trading-forum (Regulatory Concerns; chunk ae06ef0f75db5bec, chars 11293-13280)
| Stage | Iter | Model | Batch | Input | Output | Cache write | Cache read | Cost |
|---|---|---|---|---|---|---|---|---|
| plan | 0 | claude-opus-5 | no | 301 | 5029 | 0 | 0 | $0.1272 |
| gap | 0 | claude-sonnet-5 | no | 4798 | 1306 | 0 | 0 | $0.0227 |
| gap | 1 | claude-sonnet-5 | no | 16814 | 2000 | 0 | 0 | $0.0536 |
| gap | 1 | claude-sonnet-5 | no | 16882 | 947 | 0 | 0 | $0.0432 |
| synthesize | 0 | claude-opus-5 | no | 265 | 3146 | 57815 | 0 | $0.6581 |
| synthesize | 0 | claude-opus-5 | no | 366 | 7928 | 0 | 57815 | $0.2289 |
| synthesize | 0 | claude-opus-5 | no | 346 | 6513 | 0 | 57815 | $0.1935 |
| synthesize | 0 | claude-opus-5 | no | 386 | 8192 | 0 | 57815 | $0.2356 |
| synthesize | 0 | claude-opus-5 | no | 344 | 7997 | 0 | 57815 | $0.2306 |
| synthesize | 0 | claude-opus-5 | no | 356 | 8192 | 0 | 57815 | $0.2355 |
| synthesize | 0 | claude-opus-5 | no | 335 | 7996 | 0 | 57815 | $0.2305 |
| synthesize | 0 | claude-opus-5 | no | 360 | 8192 | 0 | 57815 | $0.2355 |
| synthesize | 0 | claude-opus-5 | no | 279 | 7577 | 0 | 57815 | $0.2197 |
| synthesize | 0 | claude-opus-5 | no | 265 | 5589 | 0 | 57815 | $0.1700 |
Total: $2.8846 (cap $3.50)
Per iteration: iteration 0: $2.7878, iteration 1: $0.0969
{
"finished_at": "2026-08-04T21:37:21.545378Z",
"iterations": 3,
"models": {
"embedder": "BAAI/bge-m3",
"gap": "claude-sonnet-5",
"plan": "claude-opus-5",
"reranker": "BAAI/bge-reranker-v2-m3",
"synthesize": "claude-opus-5",
"triage": "Qwen3-4B-Instruct-2507-Q4_K_M.gguf"
},
"n_chunks": 594,
"n_chunks_after_dedup": 565,
"n_chunks_evidence": 404,
"n_docs_extracted": 27,
"n_queries": 86,
"n_urls_fetched": 59,
"run_id": "1497b2907a55",
"started_at": "2026-08-04T20:47:57.300911Z",
"topic": "In OTC corporate bond markets, where a dealer's quoting policy itself reshapes the flow, spreads, and liquidity it will later be trained on, can a policy and the market response it induces be learned jointly as a self-consistent fixed point, and does that fixed-point policy generalize better than one trained under the standard assumption that the data-generating process is exogenous?"
}