Skip to content

Latest commit

 

History

History
118 lines (86 loc) · 5.79 KB

File metadata and controls

118 lines (86 loc) · 5.79 KB

First time setup (for developers)

These are instructions to set up an Textcavator development server. If you are going to develop Textcavator, start by following these instructions.

First-time setup (without Docker)

Prerequisites

  • Python == 3.12
  • PostgreSQL >= 14, client, server and C libraries
  • ElasticSearch 8. To avoid a lot of errors, choose the option: install elasticsearch with .zip or .tar.gz. ES wil install everything in one folder, and not all over your machine, which happens with other options.
  • Redis. Recommended installation is installing from source
  • Node.js. See .nvmrc for the recommended version.
  • Yarn

The documentation includes a recipe for installing the prerequisites in a distrobox container.

Installation

To get an instance running, do all of the following inside an activated virtualenv:

  1. Create the file backend/ianalyzer/settings_local.py.ianalyzer/settings_local.py is included in .gitignore and thus not cloned to your machine. It can be used to customise your environment. You can leave the file empty for now.
  2. Install the requirements for both the backend and frontend:
yarn postinstall
  1. For an easy setup, locate the file config/elasticsearch.yml in your Elasticsearch directory, and set the variable xpack.security.enabled: false. Alternatively, you can leave this on its default value(true), but this requires additional settings.
  2. Set up your postgres database:
psql -f backend/create_db.sql
yarn django migrate

Note

For historical reasons, the development database is called "ianalyzer". You chan change this if you want.

  1. Make a superuser account with yarn django createsuperuser

Setup with Docker

Alternatively, you can run the application via Docker:

  1. Install Docker Desktop and start it.
  2. Make an .env file next to this README, which defines the configuration for the SQL database and Redis. An example setup could look as follows:
SQL_HOST=db
SQL_PORT=5432
SQL_USER=myuser
SQL_DATABASE=mydb
SQL_PASSWORD=mysupersecretpassword
ES_HOST=elasticsearch
CELERY_BROKER=redis://redis
DATA_DIR=where/corpus/data/is/located/on/your/machine
  1. Run docker-compose up from the directory of this README. This will pull images from the Docker registry and start containers based on these images. This will take a while to set up the first time. To stop, hit ctrl-c, run docker-compose down in another terminal, or use the Docker Desktop dashboard.
  2. If you need to reinstall libraries via pip or yarn, use docker-compose up --build.

Note: you can also call the .env file .myenv and specify this during startup: docker-compose --env-file .myenv up

Add a test corpus

These instructions will add a tiny example corpus to your environment. Use this to verify that everything is working correctly. Open the file /backend/ianalyzer/settings_local.py. Copy-paste:

CORPORA = {
    'example': 'corpora_test.basic.corpus.ExampleCorpus',
}

CORPUS_SETTINGS = {
    'example': {
        'es_index': 'example-corpus'
    }
}

Save the file and close. For the next step, PostgreSQL and Elasticsearch must be running. Run in the terminal:

yarn django loadcorpora
yarn django index example

This will save the corpus configuration in the database and index the corpus data in Elasticsearch.

Running a dev environment

  1. Start your local elasticsearch server. If you installed from .zip or .tar.gz, this can be done by running {path your your elasticsearch folder}/bin/elasticsearch
  2. Activate your python environment. Start the backend server with yarn start-back. This creates an instance of the Django server at 127.0.0.1:8000.
  3. (optional) If you want to use celery, start your local redis server by running redis-server in a separate terminal.
  4. (optional) If you want to use celery, activate your python environment. Run yarn celery worker. Celery is used for long downloads and the word cloud and ngrams visualisations.
  5. Start the frontend by running yarn start-front.

Quick check

Below are some steps to go through the application and check that it's working as expected.

  • Go to http://localhost:4200 in your browser. The home page should appear. It will tell you there are no corpora to display; this is because your test corpus is still private.
  • Click "Sign in" in the top right and sign in with the superuser account you created earlier.
  • You will return to the home page and the example corpus will appear. Click on "Explore" to go to the search page. You should see search results appear.
  • Type in "to" in the search bar and press "Search".
  • Go to the "Visualizations" tab. A bar chart will appear.
  • If you are running Celery, use the field "What do you want to visualize?" to select "Frequency of the search term". An updated bar chart appears.
  • Go to the "Download" tab and press the "Download" button. Save the CSV file.
  • In the top menu, click on your username and select "Administration" in the dropdown menu to go to the Django admin site. You will stay signed in here.

That's it!

Next steps

Now that you have a working Textcavator environment, here are some common next steps:

Configure your environment -> Django project settings / Frontend environment settings

Add an existing corpus -> Adding existing corpora

Create a new Python corpus -> Writing a corpus definition in Python

Add SAML intergration in your environment -> SAML