Active particles sustain persistent out-of-equilibrium motion by consuming energy and can self-organize into coordinated structures such as swarms. During non-cooperative foraging under partial observability, the presence of another forager can act as a proxy signal for the nearby presence of food, promoting swarming. We validate this mechanism by simulating multiple self-propelled foragers harvesting from multiple resource patches in a continuous two-dimensional space with stochastic position updates and local passive sensing. We evolve a shared policy, implemented as a continuous time recurrent neural network that controls forager velocity, using an evolutionary strategy in which samples from the policy distribution are evaluated within the same rollout. The agents learn adaptive foraging, and when resource patches are removed, they exhibit swarming through aggregation. We further find that aggregation strength is inversely related to the amount of resource stored in a forager, consistent with risk-sensitive foraging. Analysis of the learned controller in minimal test runs reveals hidden states that are sensitive to stored resource, and clamping these states to represent lower reserves accelerates aggregation. These results demonstrate that internal state modulated swarming can emerge from local sensing alone in a multi-agent patch foraging environment.