Despite the low dimensionalities of dissipative viscous fluids, reinforcement learning (RL) requires many observables in fluid control problems. This is because the observables are assumed to follow a policy-independent Markov decision process in the RL framework. By including policy parameters as arguments of a value function, we construct a consistent algorithm with partial observables. Using typical examples of fluid control around a cylinder, we show that our algorithm is more stable and efficient than the existing RL algorithms, even under a small number of observations.