Business Daily Media

Men's Weekly

.

PolyU develops novel multi-modal agent to facilitate long video understanding by AI, accelerating development of generative AI-assisted video analysis

HONG KONG SAR - Media OutReach Newswire - 10 June 2025 - While Artificial Intelligence (AI) technology is evolving rapidly, AI models still struggle with understanding long videos. A research team from The Hong Kong Polytechnic University (PolyU) has developed a novel video-language agent, VideoMind, that enables AI models to perform long video reasoning and question-answering tasks by emulating humans' way of thinking.

The VideoMind framework incorporates an innovative Chain-of-Low-Rank Adaptation (LoRA) strategy to reduce the demand for computational resources and power, advancing the application of generative AI in video analysis. The findings have been submitted to the world-leading AI conferences.

A research team led by Prof. Changwen Chen, Interim Dean of the PolyU Faculty of Computer and Mathematical Sciences and Chair Professor of Visual Computing, has developed a novel video-language agent VideoMind that allows AI models to perform long video reasoning and question-answering tasks by emulating humans’ way of thinking. The VideoMind framework incorporates an innovative Chain-of-LoRA strategy to reduce the demand for computational resources and power, advancing the application of generative AI in video analysis.
A research team led by Prof. Changwen Chen, Interim Dean of the PolyU Faculty of Computer and Mathematical Sciences and Chair Professor of Visual Computing, has developed a novel video-language agent VideoMind that allows AI models to perform long video reasoning and question-answering tasks by emulating humans’ way of thinking. The VideoMind framework incorporates an innovative Chain-of-LoRA strategy to reduce the demand for computational resources and power, advancing the application of generative AI in video analysis.

Videos, especially those longer than 15 minutes, carry information that unfolds over time, such as the sequence of events, causality, coherence and scene transitions. To understand the video content, AI models therefore need not only to identify the objects present, but also take into account how they change throughout the video. As visuals in videos occupy a large number of tokens, video understanding requires vast amounts of computing capacity and memory, making it difficult for AI models to process long videos.

Prof. Changwen CHEN, Interim Dean of the PolyU Faculty of Computer and Mathematical Sciences and Chair Professor of Visual Computing, and his team have achieved a breakthrough in research on long video reasoning by AI. In designing VideoMind, they made reference to a human-like process of video understanding, and introduced a role-based workflow. The four roles included in the framework are: the Planner, to coordinate all other roles for each query; the Grounder, to localise and retrieve relevant moments; the Verifier, to validate the information accuracy of the retrieved moments and select the most reliable one; and the Answerer, to generate the query-aware answer. This progressive approach to video understanding helps address the challenge of temporal-grounded reasoning that most AI models face.

Another core innovation of the VideoMind framework lies in its adoption of a Chain-of-LoRA strategy. LoRA is a finetuning technique emerged in recent years. It adapts AI models for specific uses without performing full-parameter retraining. The innovative chain-of-LoRA strategy pioneered by the team involves applying four lightweight LoRA adapters in a unified model, each of which is designed for calling a specific role. With this strategy, the model can dynamically activate role-specific LoRA adapters during inference via self-calling to seamlessly switch among these roles, eliminating the need and cost of deploying multiple models while enhancing the efficiency and flexibility of the single model.

VideoMind is open source on GitHub and Huggingface. Details of the experiments conducted to evaluate its effectiveness in temporal-grounded video understanding across 14 diverse benchmarks are also available. Comparing VideoMind with some state-of-the-art AI models, including GPT-4o and Gemini 1.5 Pro, the researchers found that the grounding accuracy of VideoMind outperformed all competitors in challenging tasks involving videos with an average duration of 27 minutes. Notably, the team included two versions of VideoMind in the experiments: one with a smaller, 2 billion (2B) parameter model, and another with a bigger, 7 billion (7B) parameter model. The results showed that, even at the 2B size, VideoMind still yielded performance comparable with many of the other 7B size models.

Prof. Chen said, "Humans switch among different thinking modes when understanding videos: breaking down tasks, identifying relevant moments, revisiting these to confirm details and synthesising their observations into coherent answers. The process is very efficient with the human brain using only about 25 watts of power, which is about a million times lower than that of a supercomputer with equivalent computing power. Inspired by this, we designed the role-based workflow that allows AI to understand videos like human, while leveraging the chain-of-LoRA strategy to minimise the need for computing power and memory in this process."

AI is at the core of global technological development. The advancement of AI models is however constrained by insufficient computing power and excessive power consumption. Built upon a unified, open-source model Qwen2-VL and augmented with additional optimisation tools, the VideoMind framework has lowered the technological cost and the threshold for deployment, offering a feasible solution to the bottleneck of reducing power consumption in AI models.

Prof. Chen added, "VideoMind not only overcomes the performance limitations of AI models in video processing, but also serves as a modular, scalable and interpretable multimodal reasoning framework. We envision that it will expand the application of generative AI to various areas, such as intelligent surveillance, sports and entertainment video analysis, video search engines and more."


Hashtag: #PolyU #AI #LLMs #VideoAnalysis #IntelligentSurveillance #VideoSearch

The issuer is solely responsible for the content of this announcement.

News from Asia

‘War orphans’ express gratitude to Chinese foster parents

BEIJING, CHINA - Media OutReach Newswire - 21 February 2026 – Organized by the Japanese Repatriates and Japan-China Friendship Association, a delegation of 90 Japanese "war orphans," along with th...

Keeper Security Expands Relationship With Ingram Micro to Broaden Availability of Privileged Access Management in Singapore

Expansion strengthens cybersecurity resilience by delivering a modern, scalable privileged access solution SINGAPORE - Media OutReach Newswire - 23 February 2026 – Keeper Security, the leading ze...

Trad To Tech: Craftsmanship Growing Inside the Most Beautiful Homes as MIFF Leads the Way

KUALA LUMPUR, MALAYSIA - Media OutReach Newswire - 23 February 2026 - At the Malaysian International Furniture Fair (MIFF), a master craftsperson brings a solid wood tabletop to fruition, overseei...

Anaplan Launches AWS Data Center in Singapore to Enhance Global Reach and Support Local Enterprises

New location expands company’s global infrastructure, while offering faster data processing, robust security measures and regulatory compliance SINGAPORE - Media OutReach Newswire - 23 February ...

Lumen Technologies expands APAC cybersecurity capabilities in collaboration with Palo Alto Networks

SINGAPORE - Media OutReach Newswire - 23 February 2026 - Lumen Technologies has achieved the Palo Alto Networks NextWave Cortex XSIAM Select Specialisation Status in Singapore. This specialisation...

The World’s 100 Best Coffee Shops: Asia Pacific’s Notable Winners

Four Coffee Shops from Australia, Singapore and Malaysia Ranked in Top 10 SINGAPORE - Media OutReach Newswire - 23 February 2026 - The second edition of THE WORLD'S 100 BEST COFFEE SHOPS 2026 wi...

Esperanza Securities Introduces the First SFC-permitted Tokenized Investment for Live Entertainment in Asia Pacific

HONG KONG SAR - Media OutReach Newswire - 23 February 2026 - Esperanza Fintech (Securities) Limited ("Esperanza Securities", or "Company") announced today that, following the granting of the forma...

Tim Hortons® Singapore Marks Major Milestone with Official MUIS Halal Certification Ahead of the Festive Season

SINGAPORE - Media OutReach Newswire - 23 February 2026 - Tim Hortons® Singapore is pleased to announce that it has officially received Halal certification from the Majlis Ugama Islam Singapura (...

SICPA secures major European award for UK Vaping Duty Stamps Program

Swiss technology company SICPA secured a landmark traceability contract, in partnership with Spectra Systems Corporation’s subsidiary, Cartor Security Printers (Cartor), reinforcing its global lead...

Vinfast Middle East Signs MoU with PlusX Electric to Strengthen EV Ownership Experience in the UAE

DUBAI, UAE - Media OutReach Newswire - 23 February 2026 - VinFast today announced the signing of a Memorandum of Understanding (MoU) with PlusX Electric, a DEWA-approved EV charging and electric m...

Why I Decided to Build a Better Way to Build Homes

Why does building a home still feel like stepping into the unknown? In an industry where costs blow out and decisions come too late, certainty has...

Leonardo.Ai reveals new brand, expanding its creator-first platform for the next era of generative AI

The company has also launched its developer API to empower creators and builders to integrate AI into their workflows SYDNEY, Australia – 19 Febr...

Psychosocial injury risk starts inside workplace microcultures

Psychological injury is now one of the most expensive categories of workers compensation claims in Australia, with Safe Work Australia reporting t...

2025 Thryv Business and Consumer Report - Australian small businesses show grit under pressure

Australia’s small businesses are powering ahead with optimism, resilience and discipline, however, mounting pressures on costs, wellbeing and cons...

Security by Default: Why 2026 Will Force Organisations to Rethink Cloud and AI

financial accountability to how they run cloud and AI, according to leading Australian systems integrator, Brennan. Based on customer insights...

UNSW launches plan to help Aussie startups scale overseas

UNSW Launches Global Innovation Foundry to Scale 100 Australian Startups Internationally New initiative provides startups and spinouts with direc...