Skip to content

Improve Snippet function #399

Description

@NimaSarajpoor

As I was working with snippets, I noticed the three following issues:

1- Apparently the indices of the snippet start exactly at the multiplies of m, which may not be the case for some data.

2- The function provides only one profile per snippet. However, that snippet repeats throughout the whole time series as higher fraction means it has been repeated more. So, it should be useful to have that information in the output, where I can see all the profiles (the starting index) of each snippet.

3- The fraction output is not sorted. I think it is better to sort it in descending order and accordingly provide the other outputs for the same order.

Activity

  1. changed the title [-]Improving Snippet function[/-] [+]Improve Snippet function[/+] on Jun 3, 2021
  2. seanlaw commented on Jun 4, 2021

    @seanlaw
    Contributor

    @ninimama If it's okay with you, I'd like to move this over to the discussion section as I think it would be healthy/useful to discuss/debate each of the points above. We should keep in mind that STUMPY doesn't try to be everything for everyone and our goal is to reproduce the published work.

  3. NimaSarajpoor commented on Jun 4, 2021

    @NimaSarajpoor
    CollaboratorAuthor

    @seanlaw

    Sure. I think it is good to move 1 & 3 to discussion. However, if you would like to reproduce the results of the paper, I think the snippet function should produce the indices of a set of subsequences similar to each snippet (e.g: check out Fig 24 of the paper or Fig 19).

    However, if you think the ultimate goal of snippets is to merely provide a set of representative sequences rather than how they appear in the time series, then we can move 2 to discussion as well.

  4. seanlaw commented on Jun 4, 2021

    @seanlaw
    Contributor

    @ninimama I haven't looked into this part much so maybe you can help me understand how the indices are calculated? It appears that it has do with finding cross-over points.

    However, if you think the ultimate goal of snippets is to merely provide a set of representative sequences rather than how they appear in the time series, then we can move 2 to discussion as well.

    I think this is what I'm trying to get at with an open discussion. I am open to being convinced :)

  5. NimaSarajpoor commented on Jun 4, 2021

    @NimaSarajpoor
    CollaboratorAuthor

    This is part of the snippet code. If I understand correctly, mask contains the indices of distance profile corresponds to the snippet (in boolean). And, by summing over that, we calculate the fraction.

    image

    So, if my understanding is correct, all we need to do is to return those indices.

    Of course there might be some overlaps between one subsequence to the next, but we don't need to worry about that. We just need to return those indices. So, coloring them should give us the Fig. 24 of the paper.

    Please feel free to close this one and move the questions to the discussion.

  6. locked and limited conversation to collaborators on Jun 4, 2021
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions