These are instructions to set up an Textcavator development server. If you are going to develop Textcavator, start by following these instructions.
- Python == 3.12
- PostgreSQL >= 14, client, server and C libraries
- ElasticSearch 8. To avoid a lot of errors, choose the option: install elasticsearch with .zip or .tar.gz. ES wil install everything in one folder, and not all over your machine, which happens with other options.
- Redis. Recommended installation is installing from source
- Node.js. See .nvmrc for the recommended version.
- Yarn
The documentation includes a recipe for installing the prerequisites in a distrobox container.
To get an instance running, do all of the following inside an activated virtualenv:
- Create the file
backend/ianalyzer/settings_local.py.ianalyzer/settings_local.pyis included in .gitignore and thus not cloned to your machine. It can be used to customise your environment. You can leave the file empty for now. - Install the requirements for both the backend and frontend:
yarn postinstall- For an easy setup, locate the file
config/elasticsearch.ymlin your Elasticsearch directory, and set the variablexpack.security.enabled: false. Alternatively, you can leave this on its default value(true), but this requires additional settings. - Set up your postgres database:
psql -f backend/create_db.sql
yarn django migrateNote
For historical reasons, the development database is called "ianalyzer". You chan change this if you want.
- Make a superuser account with
yarn django createsuperuser
Alternatively, you can run the application via Docker:
- Install Docker Desktop and start it.
- Make an .env file next to this README, which defines the configuration for the SQL database and Redis. An example setup could look as follows:
SQL_HOST=db
SQL_PORT=5432
SQL_USER=myuser
SQL_DATABASE=mydb
SQL_PASSWORD=mysupersecretpassword
ES_HOST=elasticsearch
CELERY_BROKER=redis://redis
DATA_DIR=where/corpus/data/is/located/on/your/machine
- Run
docker-compose upfrom the directory of this README. This will pull images from the Docker registry and start containers based on these images. This will take a while to set up the first time. To stop, hitctrl-c, rundocker-compose downin another terminal, or use the Docker Desktop dashboard. - If you need to reinstall libraries via pip or yarn, use
docker-compose up --build.
Note: you can also call the .env file .myenv and specify this during startup:
docker-compose --env-file .myenv up
These instructions will add a tiny example corpus to your environment. Use this to verify that everything is working correctly. Open the file /backend/ianalyzer/settings_local.py. Copy-paste:
CORPORA = {
'example': 'corpora_test.basic.corpus.ExampleCorpus',
}
CORPUS_SETTINGS = {
'example': {
'es_index': 'example-corpus'
}
}Save the file and close. For the next step, PostgreSQL and Elasticsearch must be running. Run in the terminal:
yarn django loadcorpora
yarn django index exampleThis will save the corpus configuration in the database and index the corpus data in Elasticsearch.
- Start your local elasticsearch server. If you installed from .zip or .tar.gz, this can be done by running
{path your your elasticsearch folder}/bin/elasticsearch - Activate your python environment. Start the backend server with
yarn start-back. This creates an instance of the Django server at127.0.0.1:8000. - (optional) If you want to use celery, start your local redis server by running
redis-serverin a separate terminal. - (optional) If you want to use celery, activate your python environment. Run
yarn celery worker. Celery is used for long downloads and the word cloud and ngrams visualisations. - Start the frontend by running
yarn start-front.
Below are some steps to go through the application and check that it's working as expected.
- Go to
http://localhost:4200in your browser. The home page should appear. It will tell you there are no corpora to display; this is because your test corpus is still private. - Click "Sign in" in the top right and sign in with the superuser account you created earlier.
- You will return to the home page and the example corpus will appear. Click on "Explore" to go to the search page. You should see search results appear.
- Type in "to" in the search bar and press "Search".
- Go to the "Visualizations" tab. A bar chart will appear.
- If you are running Celery, use the field "What do you want to visualize?" to select "Frequency of the search term". An updated bar chart appears.
- Go to the "Download" tab and press the "Download" button. Save the CSV file.
- In the top menu, click on your username and select "Administration" in the dropdown menu to go to the Django admin site. You will stay signed in here.
That's it!
Now that you have a working Textcavator environment, here are some common next steps:
Configure your environment -> Django project settings / Frontend environment settings
Add an existing corpus -> Adding existing corpora
Create a new Python corpus -> Writing a corpus definition in Python
Add SAML intergration in your environment -> SAML