Related Experiment Video
Updated: Oct 19, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
752
API2CAN: a dataset & service for canonical utterance generation for REST APIs.
Mohammad-Ali Yaghoub-Zadeh-Fard1, Boualem Benatallah2
1UNSW Sydney, Kensington, Australia. m.yaghoubzadehfard@unsw.edu.au.
BMC Research Notes
|September 23, 2021
Summary
This study introduces a new dataset to train models that generate user utterances for application programming interfaces (APIs). This facilitates building natural language interfaces (NLIs) by overcoming the lack of domain-independent training data.
Area of Science:
- Natural Language Processing
- Human-Computer Interaction
- Software Engineering
Background:
- Natural language interfaces (NLIs), such as chatbots, are increasingly used to interact with applications by executing underlying APIs.
- Supervised methods for building NLIs require extensive datasets of user utterances paired with corresponding APIs.
- A significant challenge is the scarcity of training data for developing domain-independent NLI models.
Purpose of the Study:
- To address the lack of training samples for domain-independent NLI models.
- To propose a novel dataset specifically designed for training supervised models to generate initial utterances for APIs.
- To facilitate the creation of more robust and versatile natural language interfaces.
Main Methods:
- A dataset of 14,370 API method-utterance pairs was automatically generated.
- API method descriptions were converted into user utterances.
- The dataset underwent manual cleaning to ensure high quality and usability.
- Accompanying microservices were developed to aid in sample collection.
Main Results:
- The creation of a large-scale, manually curated dataset for API-to-utterance generation.
- Development of supporting microservices to streamline the data collection process.
- The dataset enables training of models for generating diverse and accurate initial utterances for APIs.
Conclusions:
- The proposed dataset and accompanying tools significantly lower the barrier to entry for developing supervised natural language interfaces.
- This resource facilitates research and development in domain-independent NLI systems.
- Enables more efficient creation of high-quality training data for translating API methods into natural language utterances.
More Related Videos
Related Concept Videos
Larynx
2.5K
The human larynx, often referred to as the voice box, is an intricate organ located in the neck. It serves as a pathway for air to enter the lungs during respiration and is an essential component of voice production.
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
Anatomy of the Larynx
The larynx consists of various components, including cartilage, muscles, and vocal cords. Its structure includes three large unpaired cartilages—the thyroid, cricoid, and epiglottis—and three smaller paired cartilages—the arytenoids,...
2.5K
The Auditory Ossicles
2.2K
The auditory ossicles of the middle ear transmit sounds from the air as vibrations to the fluid-filled cochlea. The auditory ossicles consist of two malleus (hammer) bones, two incus (anvil) bones, and two stapes (stirrups), one on each side. These bones develop during the fetal stage and are the ones to ossify first. They are fully mature at birth and do not grow afterward.
The aptly named stapes look very much like a stirrup. The three ossicles are unique to mammals, and each plays a role in...
The aptly named stapes look very much like a stirrup. The three ossicles are unique to mammals, and each plays a role in...
2.2K
Soundness of Cement
283
The soundness of cement refers to the ability of cement paste to retain its volume after setting. Unsound cement can lead to expansion and structural damage due to the presence of free lime, magnesia, and calcium sulfate. Free lime hydrates very slowly, expanding and causing unsoundness, which is difficult to detect because it intercrystallizes with other compounds. Magnesia also reacts with water, forming crystals that can disrupt the cement's structure. Calcium sulfate can create...
283
Sound Intensity Level
4.4K
Humans perceive sound by hearing. The human ear helps sound waves reach the brain, which then interprets the waves and creates the perception of hearing. The loudness of the environment in which a person is located determines whether they can distinguish between different sound sources.
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
The human ear can perceive an extensive range of sound intensity, necessitating the use of the logarithmic scale to define a physical quantity—the intensity level. It is a ratio of two intensities and...
4.4K
Data Collection by Survey
7.5K
The systematic method of obtaining and analyzing accurate information of a population is called data collection. A survey is a standard method of data collection that involves collecting information from a target human population about their experience, opinion, or knowledge of a product, service, or process. The responses are recorded and interpreted. The most common survey examples are written questionnaires, face-to-face or telephonic conversations, focus groups, and electronic (e-mail or...
7.5K
Sampling Distribution
14.9K
Given simple random samples of size n from a given population with a measured characteristic such as mean, proportion, or standard deviation for each sample, the probability distribution of all the measured characteristics is called a sampling distribution. How much the statistic varies from one sample to another is known as the sampling variability of a statistic. You typically measure the sampling variability of a statistic by its standard error. The standard error of the mean is an example...
14.9K

