In this work, we study the deception of a Linear-Quadratic-Gaussian (LQG)\nagent by manipulating the cost signals. We show that a small falsification of\nthe cost parameters will only lead to a bounded change in the optimal policy.\nThe bound is linear on the amount of falsification the attacker can apply to\nthe cost parameters. We propose an attack model where the attacker aims to\nmislead the agent into learning a `nefarious' policy by intentionally\nfalsifying the cost parameters. We formulate the attack's problem as a convex\noptimization problem and develop necessary and sufficient conditions to check\nthe achievability of the attacker's goal.\n We showcase the adversarial manipulation on two types of LQG learners: the\nbatch RL learner and the other is the adaptive dynamic programming (ADP)\nlearner. Our results demonstrate that with only 2.296% of falsification on the\ncost data, the attacker misleads the batch RL into learning the 'nefarious'\npolicy that leads the vehicle to a dangerous position. The attacker can also\ngradually trick the ADP learner into learning the same `nefarious' policy by\nconsistently feeding the learner a falsified cost signal that stays close to\nthe actual cost signal. The paper aims to raise people's awareness of the\nsecurity threats faced by RL-enabled control systems.\n