Adaptive Window Diffusion Decoding for Long-Context Language Generation Under Memory Constraints
DOI:
https://doi.org/10.54097/v7jevw42Keywords:
Diffusion language models, long-context generation, memory-constrained inference, adaptive window decoding, masked denoising, context compressionAbstract
Masked diffusion language models (MDLMs) are attractive for long-form generation because they denoise all positions with bidirectional context, but this advantage also creates a severe inference-time memory bottleneck: each denoising step requires dense attention over the active sequence. Recent sliding-window and speculative wrappers have shown that the bottleneck can be reduced without retraining the base model. However, fixed window sizes remain inefficient under heterogeneous documents, where some regions require broad context and others can be decoded safely with smaller memory footprints. This paper proposes Adaptive Window Diffusion Decoding (AWDD), a memory-budgeted inference framework that adjusts window size, overlap commitment, and context-summary capacity according to observable uncertainty in the partially masked sequence. AWDD uses a lightweight dependency-pressure score, a bounded summary buffer, and confidence-aware overlap commitment to preserve long-range terms while keeping peak memory below a fixed budget. We derive the memory and computational bounds of the scheduler and provide a reproducible CPU benchmark rather than unsupported large-model claims. On a controlled long-range dependency benchmark and a real-text sanity set extracted from the uploaded SW-SpeedDLM article, AWDD improves long-range entity reconstruction over fixed-window baselines while using substantially less peak memory than a large fixed window. These results support adaptive windowing as a practical direction for memory-constrained diffusion decoding and provide code, raw results, and document-level figures for replication.
Downloads
References
[1] Teng, D., Rhee, M., Qin, Y., Zi, B., & Liu, W. (2026). SW-SpeedDLM: Sliding window speculative decoding for diffusion language models under long context constraints. Mathematics, 14(12), 2137. https://doi.org/10.3390/math14122137
[2] Sahoo, S. S., Arriola, M., Schiff, Y., Gokaslan, A., Marroquin, E., Chiu, J. T., Rush, A., & Kuleshov, V. (2024). Simple and effective masked diffusion language models. arXiv:2406.07524.
[3] Nie, S., Zhu, F., You, Z., Zhang, X., Ou, J., Hu, J., Zhou, J., Lin, Y., Wen, J.-R., & Li, C. (2025). Large language diffusion models. arXiv:2502.09992.
[4] Austin, J., Johnson, D. D., Ho, J., Tarlow, D., & van den Berg, R. (2021). Structured denoising diffusion models in discrete state-spaces. In Advances in Neural Information Processing Systems (Vol. 34, pp. 17981–17993).
[5] Hoogeboom, E., Nielsen, D., Jaini, P., Forré, P., & Welling, M. (2021). Argmax flows and multinomial diffusion: Learning categorical distributions. arXiv:2102.05379.
[6] Lou, A., Meng, C., & Ermon, S. (2024). Discrete diffusion modeling by estimating the ratios of the data distribution. arXiv:2310.16834.
[7] Shi, J., Han, K., Wang, Z., Doucet, A., & Titsias, M. K. (2024). Simplified and generalized masked diffusion for discrete data. arXiv:2409.02908.
[8] Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Vol. 1, pp. 4171–4186). Association for Computational Linguistics. https://aclanthology.org/N19-1423
[9] Beltagy, I., Peters, M. E., & Cohan, A. (2020). Longformer: The long-document transformer. arXiv:2004.05150.
[10] Zaheer, M., et al. (2020). Big Bird: Transformers for longer sequences. In Advances in Neural Information Processing Systems (Vol. 33, pp. 17283–17297).
[11] Xiao, G., Tian, Y., Chen, B., Han, S., & Lewis, M. (2023). Efficient streaming language models with attention sinks. arXiv:2309.17453.
[12] Zhang, Z., et al. (2023). H2O: Heavy-hitter oracle for efficient generative inference of large language models. In Advances in Neural Information Processing Systems (Vol. 36).
[13] Li, Y., Huang, Y., Yang, B., Vber, B., Hashemi, H., Shu, Y., & El-Khamy, M. (2024). SnapKV: LLM knows what you are looking for before generation. arXiv:2404.14469.
[14] Cai, Z., et al. (2024). PyramidKV: Dynamic KV cache compression based on pyramidal information funneling. arXiv:2406.02069.
[15] Chevalier, A., Wettig, A., Ajith, A., & Chen, D. (2023). Adapting language models to compress contexts. arXiv:2305.02897.
[16] Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning (Vol. 202, pp. 19274–19286). PMLR. https://proceedings.mlr.press/v202/leviathan23a.html
[17] Chen, C., Borgeaud, S., Irving, G., Lespiau, J.-B., Sifre, L., & Jumper, J. (2023). Accelerating large language model decoding with speculative sampling. arXiv:2302.01318.
[18] Teng, D. (2025). PACO: Predictive auto-configuration for SLO-constrained large language model inference serving. Innovation and Technology Studies.
[19] Fu, Y., Bailis, P., Stoica, I., & Zhang, H. (2024). Break the sequential dependency of LLM inference using lookahead decoding. arXiv:2402.02057.
[20] Katharopoulos, A., Vyas, A., Pappas, N., & Fleuret, F. (2020). Transformers are RNNs: Fast autoregressive transformers with linear attention. In Proceedings of the 37th International Conference on Machine Learning (pp. 5156–5165). PMLR.
[21] Choromanski, K., et al. (2021). Rethinking attention with performers. In Proceedings of the International Conference on Learning Representations.
[22] Wang, Z., Yang, J. S., Shang, W., & Ding, J. (2026). FairPromote: Explainable and fairness-aware talent promotion prediction via adversarial debiasing and SHAP-based interpretation. IEEE Access, 14, 72890–72904. https://doi.org/10.1109/ACCESS.2026.3583411
[23] Zhang, F., Guo, Z., Ding, J., Yang, J., & Liu, W. (2026). Adaptive sensor fusion for robust perception in dense fog: A gated vision and LiDAR integration framework. Sensors, 26(12), 3728. https://doi.org/10.3390/s26123728
[24] Zi, B. (2024). Large language models for enterprise workflow automation in financial operations. Innovation and Technology Studies, 1(1), 24–29.
[25] Ding, J., Shen, Z., & Liu, W. (2026). Game-theoretic cost-sensitive adversarial training for robust cloud intrusion detection against GAN-based evasion attacks. Applied Sciences, 16(8), 3944. https://doi.org/10.3390/app16083944
[26] Zi, B. (2024). Cloud-native distributed systems for real-time payment intelligence. AI and Data Science Journal, 1(1), 51–56.
[27] Wang, B., Wang, Z., Zhao, W., Zhang, F., & Shang, W. (2026). DRL-Adapt: Deep reinforcement learning for adaptive routing convergence optimization in large-scale networks. IEEE Open Journal of the Computer Society, 7, 261–270. https://doi.org/10.1109/OJCS.2026.3612745
[28] Zhang, H. (2025). Physics-informed neural networks for high-fidelity electromagnetic field approximation in VLSI and RF EDA applications. Journal of Computing and Electronic Information Management, 18(2), 38–46.
[29] Jiao, Y., Fan, H., Yue, X., Ping, W., Sun, T., & Wang, J. (2026). Dynamic heterogeneous graph contrastive learning for uncovering collusive financial fraud. Scientific Reports, 16, 12907. https://doi.org/10.1038/s41598-026-96214-3
[30] Ping, W., Jiao, Y., Fan, H., & Zhang, X. (2026). Multimodal fraud detection in financial statements: A trimodal attention network with contrastive evidence chain construction. IEEE Access, 14, 74345–74356. https://doi.org/10.1109/ACCESS.2026.3591246
Downloads
Published
Issue
Section
License
Copyright (c) 2026 International Journal of Advanced Engineering and Technology Research

This work is licensed under a Creative Commons Attribution-NonCommercial 4.0 International License.










