Deep Reinforcement Learning for MC-MIMO-NOMA: A Corrected Benchmark Reveals No Gains over Classical Beamforming and Power Control
Keywords:
MIMO-NOMA, beamforming, power control, user scheduling, reinforcement learning, benchmarking, SLNR, WMMSE.Abstract
Deep reinforcement learning (DRL) is increasingly proposed for joint beamforming, power control and user scheduling in multi-cell multiple-input multiple-output non-orthogonal multiple access (MC-MIMO-NOMA) systems, frequently with reported gains over classical optimization. This paper is a reproducibility study and corrected benchmark. Beginning from a public DRL framework for this problem, we (i) identify and fix simulation and modelling defects that render its results uninterpretable—most critically an unknown-interference power set roughly 40 dB above the received signal, which makes the per-user signal-to-interference-plus-noise-ratio (SINR) constraint infeasible in every tested configuration, an energy-detector receiver in place of a linear combiner, an inverted successive-interference-cancellation (SIC) term, a non-functional reconfigurable-intelligent-surface model, and an inconsistent policy log-likelihood; (ii) show, on the corrected and QoS-feasible environment, that end-to-end DRL fails to learn transmit-power minimization—learned power is essentially flat in the QoS threshold even in the interference-free single-user regime, whereas the analytically optimal power varies by roughly 20 dB, and this persists when training is extended by 6.7×; and (iii) study a hybrid that pairs beamforming directions with a classical optimal power-control inner loop (the Foschini–Miljanic fixed point) and greedy admission. The hybrid recovers the expected power–QoS–fairness tradeoffs and guarantees QoS for all served users, but DRL-learned directions underperform every classical baseline we test— maximum ratio transmission (MRT), regularized zero-forcing (RZF), weighted MMSE (WMMSE), and signal-to-leakage-plus-noise-ratio (SLNR) beamforming—serving as few as one third of the users SLNR serves at the same QoS, across interference and load sweeps and over multiple seeds. Our results do not support the prevailing claim that DRL outperforms classical methods for this problem. We release a corrected, QoS-feasible, statistically rigorous benchmark and delineate the conditions under which learning could plausibly add value.





