Related Experiment Videos
Evaluation of a Hybrid Neural-Polynomial Deep Q-Network for Switching-Aware Spectrum Selection in a Controlled
Yuxuan Pan1, Ying Yan1,2, Dingyi Sun3
1Reading Academy, Nanjing University of Information Science and Technology, Nanjing 210044, China.
Abstract:
Switching-aware spectrum selection requires balancing interference avoidance against retuning costs. This paper evaluates a hybrid neural-polynomial deep Q-network (HNP-DQN) in an eight-channel, measurement-driven 5 GHz testbed with exogenous sweeping and random interference. The architecture combines learned latent features with an element-wise second-order expansion, layer normalization (LayerNorm), and a dueling Double Deep Q-Network (DDQN) backbone. Files are separated before window construction, and the evaluation includes a pre-inspected pilot and two outcome-uninspected distance and power transfers. The comparison includes strong deterministic rules and a capacity-matched multilayer perceptron (MLP) trained with the same DDQN procedure. All methods are evaluated on common trajectories with paired inference across 10 independently trained seeds and Holm correction. Under deterministic sweeping, the schedule-aware rule was given the declared initial phase and one-channel-per-step direction and used an internal step counter to track the deterministic progression. These schedule variables and the corresponding step index are absent from HNP-DQN's 48-dimensional observation. The rule matched the trajectory-wise dynamic-programming upper bound and outperformed HNP-DQN in all three settings (differences calculated as HNP-DQN minus the rule: -1.91, -3.51, and -1.97; all adjusted p=0.023). This information-asymmetric operational comparison shows the advantage attainable when the declared sweep specification and its progression are directly exploited; however, it does not provide a matched-information architectural ranking. By contrast, under random jamming, all comparisons between HNP-DQN and the threshold rule were inconclusive after correction, as were all six capacity-matched comparisons between HNP-DQN and the MLP. Of the 24 exploratory ablation comparisons, one favored the γ=0 variant in the random pilot, whereas the other 23 were inconclusive. Accordingly, this paper provides an information-aware, approximately parameter-matched reference for evaluating when explicit knowledge, observation-based control, or additional learning complexity is justified within the declared measurement-replay scope.