# Does Prodigy allow loading all files from a filepath

**URL:** <https://support.prodi.gy/t/does-prodigy-allow-loading-all-files-from-a-filepath/379>\
**Category:** Uncategorized\
**Tags:** usage, solved\
**Created:** [March 9, 2018, 11:55am UTC](https://support.prodi.gy/t/does-prodigy-allow-loading-all-files-from-a-filepath/379 "2018-03-09T11:55:05Z")\
**Posts on this page:** 1\
**Showing post:** 2

<div class="post-metadata">

**Author:** ![ines](https://sea2.discourse-cdn.com/flex020/user_avatar/support.prodi.gy/ines/32/3_2.png) [@ines](https://support.prodi.gy/u/ines)\
**Post date:** [March 9, 2018, 12:26pm UTC](https://support.prodi.gy/t/does-prodigy-allow-loading-all-files-from-a-filepath/379/2 "2018-03-09T12:26:12Z")

</div>

Out-of-the-box, Prodigy currently supports loading in data from single files of [various types](https://prodi.gy/docs/#files) – for text, that’s `.jsonl`, `.json`, `.txt` and `.csv`. You can specify the loader via the `--loader` argument on the command line. If no loader is set, Prodigy will use the file extension to pick the respective loader.

```bash
prodigy ner.teach your_dataset en_core_web_sm /path/to/data.txt

```

So if you have multiple `.txt` files and want to use them all, the easiest way would be to combine them into one file. Alternatively, you can also always write your own loader script.

If no `source` argument (file path etc.) is set on the command line, it will default to `sys.stdin`. This lets you pipe data forward from a different process, like a custom script. For example:

```bash
python load_data.py | prodigy ner.teach your_dataset en_core_web_sm

```

All your custom loader script needs to do is load the data somehow, create annotation tasks in Prodigy’s format (a dictionary with a `"text"` key) and print the dumped JSON. For example:

```python
# load_data.py
from pathlib import Path
import json

data_path = Path('/path/to/directory')
for file_path in data_path.iterdir(): # iterate over directory
    lines = Path(file_path).open('r', encoding='utf8') # open file
    for line in lines:
       task = {'text': line} # create one task for each line of text
       print(json.dumps(task)) # dump and print the JSON

```

This approach works for any file format and data type – for example, you could also load in data from a different database or via an API. If you can load your data in Python, you can use it with Prodigy 😊

There’s currently also an [open feature request](https://support.prodi.gy/t/feature-request-directories-of-text-files-as-a-source-format/186) for allowing paths to directories instead. If that’s something you’re interested in having Prodigy support out-of-the-box, you can vote for it on that thread.

---

_[View the full topic](https://support.prodi.gy/t/does-prodigy-allow-loading-all-files-from-a-filepath/379)._
