Current spatio-temporal action detection methods lack sufficient capabilities in extracting and comprehending spatio-temporal information. This paper introduces an end-to-end Adaptive Cross-Scale Fusion Encoder-Decode...Current spatio-temporal action detection methods lack sufficient capabilities in extracting and comprehending spatio-temporal information. This paper introduces an end-to-end Adaptive Cross-Scale Fusion Encoder-Decoder (ACSF-ED) network to predict the action and locate the object efficiently. In the Adaptive Cross-Scale Fusion Spatio-Temporal Encoder (ACSF ST-Encoder), the Asymptotic Cross-scale Feature-fusion Module (ACCFM) is designed to address the issue of information degradation caused by the propagation of high-level semantic information, thereby extracting high-quality multi-scale features to provide superior features for subsequent spatio-temporal information modeling. Within the Shared-Head Decoder structure, a shared classification and regression detection head is constructed. A multi-constraint loss function composed of one-to-one, one-to-many, and contrastive denoising losses is designed to address the problem of insufficient constraint force in predicting results with traditional methods. This loss function enhances the accuracy of model classification predictions and improves the proximity of regression position predictions to ground truth objects. The proposed method model is evaluated on the popular dataset UCF101-24 and JHMDB-21. Experimental results demonstrate that the proposed method achieves an accuracy of 81.52% on the Frame-mAP metric, surpassing current existing methods.展开更多
人工智能(artificial intelligence,AI)的快速发展,正在重塑工程结构设计的研究图景.在力学长期奠定的理论框架之上,大语言模型(large language model,LLM)等新一代AI技术,为结构设计提供了新的认知工具与方法视角.随着结构形态空间的...人工智能(artificial intelligence,AI)的快速发展,正在重塑工程结构设计的研究图景.在力学长期奠定的理论框架之上,大语言模型(large language model,LLM)等新一代AI技术,为结构设计提供了新的认知工具与方法视角.随着结构形态空间的持续拓展,以及多尺度与多物理场耦合问题的不断涌现,依赖经验积累与局部试探的传统设计路径正逐渐逼近复杂性边界.在这样的背景下,AI不仅为高维设计空间探索、知识表达与跨学科融合提供了新的可能,也在拓展人类理解复杂系统与开展决策分析的方式.面向未来,结构设计有望在AI推动下形成“设计-制造-测试”闭环体系、多物理场统一设计范式和智能结构系统等新的发展路径,从而为结构科学在新的时代背景下持续拓展认知边界提供重要契机.展开更多
基金support for this work was supported by Key Lab of Intelligent and Green Flexographic Printing under Grant ZBKT202301.
摘要Current spatio-temporal action detection methods lack sufficient capabilities in extracting and comprehending spatio-temporal information. This paper introduces an end-to-end Adaptive Cross-Scale Fusion Encoder-Decoder (ACSF-ED) network to predict the action and locate the object efficiently. In the Adaptive Cross-Scale Fusion Spatio-Temporal Encoder (ACSF ST-Encoder), the Asymptotic Cross-scale Feature-fusion Module (ACCFM) is designed to address the issue of information degradation caused by the propagation of high-level semantic information, thereby extracting high-quality multi-scale features to provide superior features for subsequent spatio-temporal information modeling. Within the Shared-Head Decoder structure, a shared classification and regression detection head is constructed. A multi-constraint loss function composed of one-to-one, one-to-many, and contrastive denoising losses is designed to address the problem of insufficient constraint force in predicting results with traditional methods. This loss function enhances the accuracy of model classification predictions and improves the proximity of regression position predictions to ground truth objects. The proposed method model is evaluated on the popular dataset UCF101-24 and JHMDB-21. Experimental results demonstrate that the proposed method achieves an accuracy of 81.52% on the Frame-mAP metric, surpassing current existing methods.
摘要人工智能(artificial intelligence,AI)的快速发展,正在重塑工程结构设计的研究图景.在力学长期奠定的理论框架之上,大语言模型(large language model,LLM)等新一代AI技术,为结构设计提供了新的认知工具与方法视角.随着结构形态空间的持续拓展,以及多尺度与多物理场耦合问题的不断涌现,依赖经验积累与局部试探的传统设计路径正逐渐逼近复杂性边界.在这样的背景下,AI不仅为高维设计空间探索、知识表达与跨学科融合提供了新的可能,也在拓展人类理解复杂系统与开展决策分析的方式.面向未来,结构设计有望在AI推动下形成“设计-制造-测试”闭环体系、多物理场统一设计范式和智能结构系统等新的发展路径,从而为结构科学在新的时代背景下持续拓展认知边界提供重要契机.