Results 11 to 20 of about 1,624,755 (247)
Documentation to facilitate communication between dataset creators and consumers.
Timnit Gebru +6 more
openaire +2 more sources
Dataset: LoED: The LoRaWAN at the Edge Dataset [PDF]
Accepted for publication at The 3rd International SenSys+BuildSys Workshop on Data: Acquisition to Analysis (DATA '20)
Laksh Bhatia +3 more
openaire +3 more sources
Dataset Distillation for Medical Dataset Sharing
Accepted by AAAI-23 Workshop on Representation Learning for Responsible Human-Centric ...
Guang Li 0008 +3 more
openaire +2 more sources
The BreakingNews Dataset [PDF]
We present BreakingNews, a novel dataset with approximately 100K news articles including images, text and captions, and enriched with heterogeneous meta-data (e.g. GPS coordinates and popularity metrics). The tenuous connection between the images and text in news data is appropriate to take work at the intersection of Computer Vision and Natural ...
Ramisa, Arnau +3 more
openaire +3 more sources
A Dataset of Dockerfiles [PDF]
Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of approximately 178,000 unique Dockerfiles collected from GitHub.
Jordan Henkel +3 more
openaire +2 more sources
The phylogeny of a dataset [PDF]
ABSTRACTThe field of evolutionary biology offers many approaches to study the changes that occur between and within generations of species; these methods have recently been adopted by cultural anthropologists, linguists and archaeologists to study the evolution of physical artifacts.
Andrea K. Thomer, Nicholas M. Weber
openaire +1 more source
Despite the utility and benefits of omnidirectional images in robotics and automotive applications, there are no datasets of omnidirectional images available with semantic segmentation, depth map, and dynamic properties. This is due to the time cost and human effort required to annotate ground truth images.
Ahmed Rida Sekkat +3 more
openaire +1 more source
UniToBrain Dataset: A Brain Perfusion Dataset
The CT perfusion (CTP) is a medical exam for measuring the passage of a bolus of contrast solution through the brain on a pixel-by-pixel basis. The objective is to draw "perfusion maps" (namely cerebral blood volume, cerebral blood flow and time to peak) very rapidly for ischemic lesions, and to be able to distinguish between core and penumubra regions.
Daniele Perlo +5 more
openaire +3 more sources
Meta-Dataset: A Dataset of Datasets for Learning to Learn from Few Examples
Code available at https://github.com/google-research/meta ...
Eleni Triantafillou +10 more
openaire +4 more sources
We introduce TechQA, a domain-adaptation question answering dataset for the technical support domain. The TechQA corpus highlights two real-world issues from the automated customer support domain. First, it contains actual questions posed by users on a technical forum, rather than questions generated specifically for a competition or a task. Second, it
Vittorio Castelli +20 more
openaire +2 more sources

