Skip to content

[FR] Exclude Archived Repositories from Search Results by Default #663

Description

@philippe-granet

While configuring the tool to allow searching through archived repositories, I noticed that they are currently included by default in search results.
This creates noise in the output, as archived repositories are usually not relevant for most searches.
Archived repositories are generally inactive and rarely needed, most users will not want them in their default results.
Filtering them out improves relevance and search performance.

💡 Proposed Enhancement:

  • Change the default search behavior so that archived repositories are filtered out.
  • Add a configuration option to control the default search behavior (include or not archived repositories).

Thanks in advance for your feedback, and for all the great work on this project! 🙏

Activity

  1. brendan-kellam commented on Dec 5, 2025

    @brendan-kellam
    Contributor

    Thanks for raising. I agree that it probably makes sense to exclude archived (& forked) repositories by default and then have a setting to configure this behaviour. If you have bandwidth to contribute, happy to walk through how I think this should be implemented. If not, will try to get to it when I can.

  2. philippe-granet commented on Dec 5, 2025

    @philippe-granet
    Author

    Sorry, can't help, I'm not a Typescript developer

  3. Kushal-Nandha commented on Dec 7, 2025

    @Kushal-Nandha

    Hey @brendan-kellam. I'd be happy to take this on. I've thought of an approach, would love to discuss.

  4. brendan-kellam commented on Dec 7, 2025

    @brendan-kellam
    Contributor

    Hey @brendan-kellam. I'd be happy to take this on. I've thought of an approach, would love to discuss.

    @Kushal-Nandha hey, yea sure - what is the approach you are thinking?

  5. Kushal-Nandha commented on Dec 7, 2025

    @Kushal-Nandha

    Once the user searches for a query, before sending it to the parser, we'll check if the query contains the keyword "archived"; if not, we can modify the query to append " archived:no", else let the query be as it is.
    Similar approach for forked repos.

  6. brendan-kellam commented on Dec 12, 2025

    @brendan-kellam
    Contributor

    Yea this approach sounds good to me. In terms of implementation, I would say we should operate on the QueryIR rather than the query string. I think the QueryVisitor in ir.ts will be useful in order to determine if a archived keyword is included or not.

    Additionally, for the setting, I would put it in orgMetadataSchema

  7. Kushal-Nandha commented on Dec 15, 2025

    @Kushal-Nandha

    Great. Since we'll be operating on the QueryIR to filter out the archived repos, in the RawConfig.ts, we'll need to add the two missing constants FLAG_YES_ARCHIVED and FLAG_YES_FORKS (we'll need to add these because in the parser.ts we weren't adding any flags for archived:yes as it used to be default). We can then check the tree for fork/archived keywords.

    And can you brief a bit more on orgMetadataSchema 's setting.

    Thank you!

  8. brendan-kellam commented on Dec 16, 2025

    @brendan-kellam
    Contributor

    I see what you're saying, however, RawConfig.ts is a generated file from the protobuf file query.proto, which defines zoekt's gRPC api surface. I would prefer to not modify zoekt's api surface if possible.

    Some context - the "life of a query" is as follows:

    1. we use Lezer to parse a query string into a abstract syntax tree (AST), let's call this the Lezer AST.
    2. we then transform the Lezer AST into a different syntax tree the zoekt gRPC api expects. Let's call this the zoekt AST. This is also our current intermediate representation (IR), QueryIR.
    3. finally, we construct the actual search request we send to zoekt, where we modify the query a bit and add additional branch and repo_set expressions to the query.

    When transforming from the lezer AST -> zoekt AST, we loose some information (like if the query originally contained a archived keyword). Using the Lezer AST, we could do something like the following:

    const hasArchivedKeyword = (tree: Tree): boolean => {
        const cursor = tree.cursor();
        
        do {
            if (cursor.type.id === ArchivedExpr) {
                return true;
            }
        } while (cursor.next());
        
        return false;
    }
    

    It's a bit of scope creep, but my take is that we should just use the lezer AST where possible. createZoektSearchRequest can accept a Tree and then call transformTreeToIR to convert the lezer AST into the expected zoekt AST. We can delete ir.ts entirely. Phrased another way, the lezer AST will become our IR. There are a few areas in the code where we are constructing the IR manually (e.g., fileSourceApi.ts and codeNav/api.ts). I asked GPT and doesn't seem like there is a good way of manually constructing a lezer AST directly, so we should honestly just switch these to construct a query syntax string and let the parser handle building the AST. Let me know if this needs any clarification.

    For the setting, we already have a metadata field on the Org table in the database that adhears to this schema. We are storing the anonymousAccessEnabled option in there, so I think it would be easiest to just add an addition field for includeArchivedReposByDefault and includeForkedReposByDefault (or equivalent).

  9. Kushal-Nandha commented on Dec 21, 2025

    @Kushal-Nandha

    Thank you for the explanation, Brendan. I've got it all cleared now. I've implemented the changes locally in the parser.ts, and also the schema changes to orgMetadataSchema.
    Now, there's another case, as we are defaulting the archived and fork to no, let's say a user specifically searches for an "archived" repository using the repo keyword. There will be no results displayed (as the archived:no will override the repo: filter).
    The solution: We set the two booleans (hasArchived and hasFork - the ones we define) to true for the case RepoExpr in the parser.ts. This will let the forked and archived repositories to pass through.

    Are there any other edge cases that needs to be considered?

  10. philippe-granet commented on Dec 21, 2025

    @philippe-granet
    Author

    @Kushal-Nandha I frequently run queries with a partial repository path (e.g., repo:my-gitlab-group) instead of specifying full repositories path. This searches across all GitLab projects in the group, but I want to exclude by default archived projects from the results.

  11. philippe-granet commented on Dec 21, 2025

    @philippe-granet
    Author

    In SourceGraph, it display a popup with skipped results (forked and archived repos) and where you can search again, selecting archived or forked repos:

    Image
  12. Kushal-Nandha commented on Dec 24, 2025

    @Kushal-Nandha

    Thank you for the inputs Philippe, really helps.

    Would you like to discuss further, @brendan-kellam? Should we go the SourceGraph way or any other approach that you can think of?
    There can be a couple of scenarios:

    • User searches for a partial path
    • User searches for a complete repo path
      For the first one, it makes sense to exclude the archived and forked repos, but not for the latter.
  13. brendan-kellam commented on Dec 26, 2025

    @brendan-kellam
    Contributor

    Thanks @philippe-granet - I like SourceGraph's approach of having a popup indicating that results were skipped, but I fear we are introducing too much scope creep & complexity here.

    I wonder if we could just simplify things to the following:

    1. Archived & forked repositories are include by default. This way we don't impact the current user experience.
    2. We add a new button to the search bar for search settings:
    Image

    https://lucide.dev/icons/settings-2

    1. Clicking on this button will reveal a dropdown with two checkboxes:
    • Exclude archived repositories
    • Exclude forked repositories
    1. These options are persisted to local storage so a) they are persistent across sessions, and b) the configuration is per-user.

    2. We plumb these options to the api by adding additional fields to the searchOptionsSchema, similar to what we are doing with isRegexEnabled and isCaseSensitivityEnabled

    @Kushal-Nandha let me know what you think.

  14. Kushal-Nandha commented on Dec 28, 2025

    @Kushal-Nandha

    Yes, Brendan, completely makes sense. I'll go ahead with this approach.
    I'll just add one bit, we'll implement this such that the archived: and fork: search keywords take precedence over the checkboxes.
    [So, if I have checked the Exclude archived repositories and I also provide the keyword archived:yes; archived repositories WILL be included in the search results.]

  15. brendan-kellam commented on Dec 28, 2025

    @brendan-kellam
    Contributor

    Yes, Brendan, completely makes sense. I'll go ahead with this approach. I'll just add one bit, we'll implement this such that the archived: and fork: search keywords take precedence over the checkboxes. [So, if I have checked the Exclude archived repositories and I also provide the keyword archived:yes; archived repositories WILL be included in the search results.]

    Hm this is a good point... I think the trouble with your proposal is that we will need to introduce the additional complexity of figuring out if the query contains a archived or fork keyword in the query AST. I think the fundamental issue is, with this change, we no longer have a single source of truth: archived/forked repos can be excluded/included by both the syntax language or the options we pass via the api. Compare this with say case sensitivity or regular expression options that can only be configured via the api (note that case sensitivity was previously configured via the query syntax via case:, but removed in #623).

    imo, we should stick to a single source of truth here, which means one of two options:

    1. We remove support for the fork and archived keywords from the query syntax. The source of truth comes from the searchOptionsSchema (e.g., includeArchivedRepos which can be yes ,no, or only).
    2. We stick to using fork: and archived: keywords in the query syntax. If Exclude archived repositories is checked, archived:no is appended. This is similar to what you initially proposed.

    My initial gut feeling is that (2) is likely easier to implement now so we should just do that. (1) is more effort and may break existing queries, but probably makes sense since it is more inline with how we are already handling case sensitivity.

    What do you think?

  16. Kushal-Nandha commented on Dec 29, 2025

    @Kushal-Nandha

    Yes, I second that @brendan-kellam . (2) will be easier to implement.

    (1) does make sense, but if we remove the support of archived and fork keywords from the syntax, then Exclude archived repositories will not remain a mere checkbox, but probably a radio/dropdown with three possible options (yes, no, only) and that imo is a little too much?

    So, if you give it a green, I'll go ahead with the (2)'nd approach.

  17. brendan-kellam commented on Dec 29, 2025

    @brendan-kellam
    Contributor

    yea let's go with 2, thanks!

  18. added and removed
    enhancementNew feature or request
    on Mar 9, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions