Présentation (2026)
Résumé
The deployment of machine learning models in high-stakes domains raises profound questions about the privacy of the data used to train them. In this talk, I will show how combinatorial and inverse optimization can provide a rigorous methodological backbone for analyzing, quantifying, and ultimately mitigating privacy risks in modern ML pipelines. I will first discuss a white-box reconstruction attack that formulates the recovery of a random forest training data as a maximum-likelihood problem solved with constraint programming. Remarkably, this approach can often reconstruct entire datasets, even from forests with only a few trees. Next, we turn to black-box access and explainability-driven interfaces. Counterfactual explanations (now increasingly needed and exposed through ML APIs) represent a powerful attack surface. Using tools from online optimization and competitive analysis, we derive tight bounds on the number of counterfactual queries required to exactly extract tree-based models and introduce new algorithms achieving provably perfect fidelity. Finally, I will examine the protection offered by differential privacy. Focusing on ε-DP random forests, we demonstrate that even models satisfying strict DP guarantees can still leak meaningful, dataset-specific information in practice, unless the privacy noise is increased to the point where the model loses most of its predictive value.
Statistiques
Total des téléchargements à partir de PolyPublie
Téléchargements par année
Provenance des téléchargements
Dimensions
