Related Experiment Video
Updated: May 14, 2025

08:25
Combining Eye-tracking Data with an Analysis of Video Content from Free-viewing a Video of a Walk in an Urban Park Environment
Published on: May 7, 2019
8.9K
Zero-shot and few-shot multimodal plastic waste classification with vision-language models
Iman Ranjbar1, Yiannis Ventikos2, Mehrdad Arashpour1
1Department of Civil Engineering, Monash University, Melbourne, VIC 3800, Australia.
Waste Management (New York, N.Y.)
|April 24, 2025
Summary
Vision-Language Models (VLMs) offer a scalable solution for recycling construction plastic waste. These models achieve high accuracy in classifying plastic types with minimal data, improving recycling efficiency.
Area of Science:
- Materials Science
- Computer Science
- Environmental Science
Background:
- The construction sector generates significant plastic waste, necessitating efficient recycling processes.
- Accurate classification of plastic waste is crucial for value retention during recycling.
- Current supervised deep learning models require extensive labeled data, limiting scalability and generalizability.
Purpose of the Study:
- To explore the application of Vision-Language Models (VLMs) for classifying construction and demolition plastic waste by resin type.
- To evaluate the efficacy of VLMs in zero-shot and few-shot learning scenarios for plastic waste classification.
- To compare VLM performance against fully supervised methods regarding accuracy, scalability, and data efficiency.
Main Methods:
- Utilized advanced VLMs for zero-shot classification of plastic waste using natural language descriptions.
- Integrated image and textual modalities within multimodal few-shot learning frameworks.
- Conducted comprehensive experiments comparing VLM performance with data-intensive, fully supervised baselines.
Main Results:
- VLMs demonstrated effective classification of end-of-life plastics with minimal to no training data.
- Achieved 70.15% accuracy in zero-shot classification using VLMs.
- Improved accuracy to 85.07% with multimodal few-shot learning, highlighting enhanced data efficiency and scalability.
Conclusions:
- VLMs offer a promising, data-efficient approach for classifying construction and demolition plastic waste.
- The zero-shot and few-shot learning capabilities of VLMs significantly enhance scalability for plastic waste management.
- VLMs present a viable alternative to traditional supervised methods, addressing challenges in data acquisition and model generalizability.
Related Concept Videos
Force Classification
1.0K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
1.0K
Classification of Systems-I
161
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
161
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Aggregates Classification
290
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
290
Classification of Systems-II
125
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
125
Vision
52.4K
Vision is the result of light being detected and transduced into neural signals by the retina of the eye. This information is then further analyzed and interpreted by the brain. First, light enters the front of the eye and is focused by the cornea and lens onto the retina—a thin sheet of neural tissue lining the back of the eye. Because of refraction through the convex lens of the eye, images are projected onto the retina upside-down and reversed.
52.4K

