Decision trees usefully represent sparse, high dimensional and noisy data.\nHaving learned a function from this data, we may want to thereafter integrate\nthe function into a larger decision-making problem, e.g., for picking the best\nchemical process catalyst. We study a large-scale, industrially-relevant\nmixed-integer nonlinear nonconvex optimization problem involving both\ngradient-boosted trees and penalty functions mitigating risk. This\nmixed-integer optimization problem with convex penalty terms broadly applies to\noptimizing pre-trained regression tree models. Decision makers may wish to\noptimize discrete models to repurpose legacy predictive models, or they may\nwish to optimize a discrete model that particularly well-represents a data set.\nWe develop several heuristic methods to find feasible solutions, and an exact,\nbranch-and-bound algorithm leveraging structural properties of the\ngradient-boosted trees and penalty functions. We computationally test our\nmethods on concrete mixture design instance and a chemical catalysis industrial\ninstance.\n