Learning directed acyclic graphs (DAGs) from data is a challenging task both\nin theory and in practice, because the number of possible DAGs scales\nsuperexponentially with the number of nodes. In this paper, we study the\nproblem of learning an optimal DAG from continuous observational data. We cast\nthis problem in the form of a mathematical programming model which can\nnaturally incorporate a super-structure in order to reduce the set of possible\ncandidate DAGs. We use the penalized negative log-likelihood score function\nwith both $\\ell_0$ and $\\ell_1$ regularizations and propose a new mixed-integer\nquadratic optimization (MIQO) model, referred to as a layered network (LN)\nformulation. The LN formulation is a compact model, which enjoys as tight an\noptimal continuous relaxation value as the stronger but larger formulations\nunder a mild condition. Computational results indicate that the proposed\nformulation outperforms existing mathematical formulations and scales better\nthan available algorithms that can solve the same problem with only $\\ell_1$\nregularization. In particular, the LN formulation clearly outperforms existing\nmethods in terms of computational time needed to find an optimal DAG in the\npresence of a sparse super-structure.\n