Modern 5G/O-RAN networks must route heterogeneous services with distinct QoS constraints (e.g., bandwidth, compute resources, and strict end-to-end latency) over shared resources. Centralized reinforcement learning (RL) controllers and synchronized multi-agent training can encounter serialization bottlenecks and delays due to stragglers. We propose Asynchronous Multi-Agent Reinforcement Learning (AMARL): one Proximal Policy Optimization (PPO) agent per service plans routes in parallel on local environment snapshots and commits resource deltas to a shared global state via a lock-guarded commit/abort mechanism that preserves feasibility. On an O-RAN-like simulator driven by 24-hour Montreal traffic traces, AMARL matches a strong single-agent PPO baseline in Grade of Service (GoS) and latency, while reducing training wall-clock by 29.9% and evaluation time by 15.0%.