Publications
Publications in reversed chronological order
2026
- Strategic Decision Focused LearningTinashe Handina, Yuehan Diao , Adam Wierman , and 1 more authorIn 40th Conference on Neural Information Processing Systems (NeurIPS 2026) , 2026
Machine learning (ML) predictions are increasingly being used to guide decision-making, giving rise to the problem of decision-focused learning (DFL) where predictors are optimized for downstream decision quality rather than accuracy alone. However, most existing work assumes a single decision-maker optimizing in isolation. This paper formalizes strategic decision-focused learning, where an ML system predicts an exogenous state that some agents observe before playing a game. For example, a park ranger may predict wildlife locations to allocate anti-poaching patrols against strategic poachers. While the exogenous state is unaffected by agent actions, predictions influence agents’ strategies and the resulting equilibrium. We find that strategic considerations fundamentally change the learning problem. In particular, we show the prediction accuracy–equilibrium payoff landscape can be non-monotonic—i.e., better predictions can degrade performance. We propose algorithmic approaches to address these challenges and validate them across benchmarks in wildlife conservation and infrastructure protection. Our theory and experiments highlight the importance of accounting for strategic interactions when designing predictors.
- Leveraging Machine-Learned Advice in Strategic Interactions with No-Regret LearnersTinashe Handina, Tongxin Li , Kishan Panaganti , and 2 more authorsIn 29th International Conference on Artificial Intelligence and Statistics (AISTATS 2026) , 2026
We study how an agent in a two-player repeated game can effectively utilize potentially imperfect advice when interacting with a no-regret learner. We characterize the advice landscape by introducing a pseudo-metric to quantify the usefulness of an advice instance. We demonstrate the pseudo-metric’s applicability through two forms of advice: simulators and payoff matrix predictions. We then show how an optimizing player, equipped with correctness guarantees on the advice, could leverage simulators to compute approximate Stackelberg strategies more efficiently, reducing the interaction complexity traditionally required and illustrating the power of good advice. Finally, we extend our analysis to settings where the advice does not have any guarantee of correctness. We find that, in general, a player cannot simultaneously guarantee near Stackelberg performance when the advice is approximately accurate and a no-regret condition when the advice is inaccurate. We do show, however, that it is possible for an advice-aided player to weakly dominate their utility in some (coarse)-correlated equilibria.
2025
- Plan for the Worst with Advice: Advice-Augmented Robust Markov Decision ProcessesTinashe Handina, Kishan Panaganti , Eric Mazumdar , and 1 more authorIn 64th IEEE Conference on Decision and Control (CDC 2025) , 2025
We consider the integration of advice into Robust Markov Decision Processes (RMDPs). While the RMDP formulation aids in modeling ambiguity with respect to transition dynamics, it is overly conservative due to its focus on worst-case instances. To move beyond the worst-case framework, we propose an advice-augmented setting in which the decision maker has access to advice in the form of a predicted transition kernel they seek to leverage to obtain better guarantees. The decision maker in this setting cares about finding a policy that performs well for both the worst case and advice transition dynamics. Thus, we define robustness and consistency as metrics the decision maker optimizes and propose a family of optimization problems whose solutions are Pareto-optimal with respect to robustness and consistency. Under standard assumptions on the ambiguity set, the optimal solutions are deterministic, Markovian, and stationary. Given a set of Pareto-optimal policies, we then provide a policy selection algorithm that achieves max-min optimality across robustness and consistency.
2024
- Understanding Model Selection For Learning In Strategic EnvironmentsTinashe Handina, and Eric MazumdarIn 38th Conference on Neural Information Processing Systems (NeurIPS 2024) , 2024
The deployment of ever-larger machine learning models reflects a growing consensus that the more expressive the model class one optimizes over–and the more data one has access to–the more one can improve performance. As models get deployed in a variety of real-world scenarios, they inevitably face strategic environments. In this work, we consider the natural question of how the interplay of models and strategic interactions affects the relationship between performance at equilibrium and the expressivity of model classes. We find that strategic interactions can break the conventional view–meaning that performance does not necessarily monotonically improve as model classes get larger or more expressive (even with infinite data). We show the implications of this result in several contexts including strategic regression, strategic classification, and multi-agent reinforcement learning. In particular, we show that each of these settings admits a Braess’ paradox-like phenomenon in which optimizing over less expressive model classes allows one to achieve strictly better equilibrium outcomes. Motivated by these examples, we then propose a new paradigm for model selection in games wherein an agent seeks to choose amongst different model classes to use as their action set in a game.
- Safe Exploitative Play with Untrusted Type BeliefsTongxin Li , Tinashe Handina, Shaolei Ren , and 1 more authorIn 38th Conference on Neural Information Processing Systems (NeurIPS 2024) , 2024
The combination of the Bayesian game and learning has a rich history, with the idea of controlling a single agent in a system composed of multiple agents with unknown behaviors given a set of types, each specifying a possible behavior for the other agents. The idea is to plan an agent’s own actions with respect to those types which it believes are most likely to maximize the payoff. However, the type beliefs are often learned from past actions and likely to be incorrect. With this perspective in mind, we consider an agent in a game with type predictions of other components, and investigate the impact of incorrect beliefs to the agent’s payoff. In particular, we formally define a tradeoff between risk and opportunity by comparing the payoff obtained against the optimal payoff, which is represented by a gap caused by trusting or distrusting the learned beliefs. Our main results characterize the tradeoff by establishing upper and lower bounds on the Pareto front for both normal-form and stochastic Bayesian games, with numerical results provided.
2022
- Chasing Convex Bodies and Functions with Black-Box AdviceNicolas Christianson , Tinashe Handina, and Adam WiermanIn Proceedings of Thirty Fifth Conference on Learning Theory , 2022
We consider the problem of convex function chasing with black-box advice, where an online decision-maker aims to minimize the total cost of making and switching between decisions in a normed vector space, aided by black-box advice such as the decisions of a machine-learned algorithm. The decision-maker seeks cost comparable to the advice when it performs well, known as consistency, while also ensuring worst-case robustness even when the advice is adversarial. We first consider the common paradigm of algorithms that switch between the decisions of the advice and a competitive algorithm, showing that no algorithm in this class can improve upon 3-consistency while staying robust. We then propose two novel algorithms that bypass this limitation by exploiting the problem’s convexity. The first, Interp, achieves (√2+ε)-consistency and O(C/ε²)-robustness for any ε > 0, where C is the competitive ratio of an algorithm for convex function chasing or a subclass thereof. The second, BdInterp, achieves (1+ε)-consistency and O(CD/ε)-robustness when the problem has bounded diameter D. Further, we show that BdInterp achieves near-optimal consistency-robustness trade-off for the special case where cost functions are α-polyhedral.
2021
- Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?Vikash Sehwag , Saeed Mahloujifar , Tinashe Handina, and 4 more authorsIn International Conference on Learning Representations , 2021
While additional training data improves the robustness of deep neural networks against adversarial examples, it presents the challenge of curating a large number of specific real-world samples. We circumvent this challenge by using additional data from proxy distributions learned by advanced generative models. We first seek to formally understand the transfer of robustness from classifiers trained on proxy distributions to the real data distribution. We prove that the difference between the robustness of a classifier on the two distributions is upper bounded by the conditional Wasserstein distance between them. Next we use proxy distributions to significantly improve the performance of adversarial training on five different datasets. For example, we improve robust accuracy by up to 7.5% and 6.7% in ℓ∞ and ℓ₂ threat model over baselines that are not using proxy distributions on the CIFAR-10 dataset. We also improve certified robust accuracy by 7.6% on the CIFAR-10 dataset. We further demonstrate that different generative models bring a disparate improvement in the performance in robust training. We propose a robust discrimination approach to characterize the impact of individual generative models and further provide a deeper understanding of why current state-of-the-art in diffusion-based generative models are a better choice for proxy distribution than generative adversarial networks.
- A Random Walk in Extensive Form Games: An Investigation into Information, Strategy-Proofness, and CredibilityTinashe Handina2021
We focus on extensive form games and look at how notions of credibility and strategy-proofness are affected when participants have varying levels of information. More precisely, we consider cases where participants have more and less information and we look at the implications for strategy-proofness and credibility. We also consider the auction setting and illustrate an auction, which is not the ascending price auction but which is ex-post Nash and credible. This shows how the ascending price auction is not the only ex-post Nash and credible auction. Lastly, we consider extensive form games with singleton information sets and show how if such games are strategy-proof, they are also ‘winner-pooling’.