Repository navigation
[FR] Exclude Archived Repositories from Search Results by Default #663
Description
Activity
brendan-kellam commented
on Dec 5, 2025 ContributorMore actionsThanks for raising. I agree that it probably makes sense to exclude archived (& forked) repositories by default and then have a setting to configure this behaviour. If you have bandwidth to contribute, happy to walk through how I think this should be implemented. If not, will try to get to it when I can.
Sorry, can't help, I'm not a Typescript developer
Hey @brendan-kellam. I'd be happy to take this on. I've thought of an approach, would love to discuss.
Hey @brendan-kellam. I'd be happy to take this on. I've thought of an approach, would love to discuss.
@Kushal-Nandha hey, yea sure - what is the approach you are thinking?
Once the user searches for a query, before sending it to the parser, we'll check if the query contains the keyword "archived"; if not, we can modify the query to append " archived:no", else let the query be as it is.
Similar approach for forked repos.brendan-kellam commented
on Dec 12, 2025 ContributorMore actionsYea this approach sounds good to me. In terms of implementation, I would say we should operate on the QueryIR rather than the query string. I think the
QueryVisitorinir.tswill be useful in order to determine if aarchivedkeyword is included or not.Additionally, for the setting, I would put it in
orgMetadataSchemaGreat. Since we'll be operating on the QueryIR to filter out the
archivedrepos, in theRawConfig.ts, we'll need to add the two missing constantsFLAG_YES_ARCHIVEDandFLAG_YES_FORKS(we'll need to add these because in theparser.tswe weren't adding any flags forarchived:yesas it used to be default). We can then check the tree for fork/archived keywords.And can you brief a bit more on
orgMetadataSchema's setting.Thank you!
brendan-kellam commented
on Dec 16, 2025 ContributorMore actionsI see what you're saying, however,
RawConfig.tsis a generated file from the protobuf file query.proto, which defines zoekt's gRPC api surface. I would prefer to not modify zoekt's api surface if possible.Some context - the "life of a query" is as follows:
- we use Lezer to parse a query string into a abstract syntax tree (AST), let's call this the Lezer AST.
- we then transform the Lezer AST into a different syntax tree the zoekt gRPC api expects. Let's call this the zoekt AST. This is also our current intermediate representation (IR),
QueryIR. - finally, we construct the actual search request we send to zoekt, where we modify the query a bit and add additional
branchandrepo_setexpressions to the query.
When transforming from the lezer AST -> zoekt AST, we loose some information (like if the query originally contained a archived keyword). Using the Lezer AST, we could do something like the following:
const hasArchivedKeyword = (tree: Tree): boolean => { const cursor = tree.cursor(); do { if (cursor.type.id === ArchivedExpr) { return true; } } while (cursor.next()); return false; }It's a bit of scope creep, but my take is that we should just use the lezer AST where possible.
createZoektSearchRequestcan accept aTreeand then calltransformTreeToIRto convert the lezer AST into the expected zoekt AST. We can deleteir.tsentirely. Phrased another way, the lezer AST will become our IR. There are a few areas in the code where we are constructing the IR manually (e.g.,fileSourceApi.tsandcodeNav/api.ts). I asked GPT and doesn't seem like there is a good way of manually constructing a lezer AST directly, so we should honestly just switch these to construct a query syntax string and let the parser handle building the AST. Let me know if this needs any clarification.For the setting, we already have a
metadatafield on the Org table in the database that adhears to this schema. We are storing theanonymousAccessEnabledoption in there, so I think it would be easiest to just add an addition field forincludeArchivedReposByDefaultandincludeForkedReposByDefault(or equivalent).Thank you for the explanation, Brendan. I've got it all cleared now. I've implemented the changes locally in the
parser.ts, and also the schema changes toorgMetadataSchema.
Now, there's another case, as we are defaulting thearchivedandforktono, let's say a user specifically searches for an "archived" repository using therepokeyword. There will be no results displayed (as thearchived:nowill override therepo:filter).
The solution: We set the two booleans (hasArchivedandhasFork- the ones we define) to true for the caseRepoExprin theparser.ts. This will let theforkedandarchivedrepositories to pass through.Are there any other edge cases that needs to be considered?
@Kushal-Nandha I frequently run queries with a partial repository path (e.g.,
repo:my-gitlab-group) instead of specifying full repositories path. This searches across all GitLab projects in the group, but I want to exclude by default archived projects from the results.Thank you for the inputs Philippe, really helps.
Would you like to discuss further, @brendan-kellam? Should we go the SourceGraph way or any other approach that you can think of?
There can be a couple of scenarios:- User searches for a partial path
- User searches for a complete repo path
For the first one, it makes sense to exclude the archived and forked repos, but not for the latter.
Thanks @philippe-granet - I like SourceGraph's approach of having a popup indicating that results were skipped, but I fear we are introducing too much scope creep & complexity here.
I wonder if we could just simplify things to the following:
- Archived & forked repositories are include by default. This way we don't impact the current user experience.
- We add a new button to the search bar for search settings:
https://lucide.dev/icons/settings-2
- Clicking on this button will reveal a dropdown with two checkboxes:
-
Exclude archived repositories -
Exclude forked repositories
-
These options are persisted to local storage so a) they are persistent across sessions, and b) the configuration is per-user.
-
We plumb these options to the api by adding additional fields to the searchOptionsSchema, similar to what we are doing with
isRegexEnabledandisCaseSensitivityEnabled
@Kushal-Nandha let me know what you think.
Reacted by Philippe GRANETYes, Brendan, completely makes sense. I'll go ahead with this approach.
I'll just add one bit, we'll implement this such that thearchived:andfork:search keywords take precedence over the checkboxes.
[So, if I have checked theExclude archived repositoriesand I also provide the keywordarchived:yes; archived repositories WILL be included in the search results.]Yes, Brendan, completely makes sense. I'll go ahead with this approach. I'll just add one bit, we'll implement this such that the
archived:andfork:search keywords take precedence over the checkboxes. [So, if I have checked theExclude archived repositoriesand I also provide the keywordarchived:yes; archived repositories WILL be included in the search results.]Hm this is a good point... I think the trouble with your proposal is that we will need to introduce the additional complexity of figuring out if the query contains a
archivedorforkkeyword in the query AST. I think the fundamental issue is, with this change, we no longer have a single source of truth: archived/forked repos can be excluded/included by both the syntax language or the options we pass via the api. Compare this with say case sensitivity or regular expression options that can only be configured via the api (note that case sensitivity was previously configured via the query syntax viacase:, but removed in #623).imo, we should stick to a single source of truth here, which means one of two options:
- We remove support for the
forkandarchivedkeywords from the query syntax. The source of truth comes from thesearchOptionsSchema(e.g.,includeArchivedReposwhich can beyes,no, oronly). - We stick to using
fork:andarchived:keywords in the query syntax. IfExclude archived repositoriesis checked,archived:nois appended. This is similar to what you initially proposed.
My initial gut feeling is that (2) is likely easier to implement now so we should just do that. (1) is more effort and may break existing queries, but probably makes sense since it is more inline with how we are already handling case sensitivity.
What do you think?
- We remove support for the
Yes, I second that @brendan-kellam . (2) will be easier to implement.
(1) does make sense, but if we remove the support of
archivedandforkkeywords from the syntax, thenExclude archived repositorieswill not remain a mere checkbox, but probably a radio/dropdown with three possible options (yes,no,only) and that imo is a little too much?So, if you give it a green, I'll go ahead with the (2)'nd approach.
brendan-kellam commented
on Dec 29, 2025 ContributorMore actionsyea let's go with 2, thanks!
- added and removedenhancementNew feature or requestNew feature or request
on Mar 9, 2026

While configuring the tool to allow searching through archived repositories, I noticed that they are currently included by default in search results.
This creates noise in the output, as archived repositories are usually not relevant for most searches.
Archived repositories are generally inactive and rarely needed, most users will not want them in their default results.
Filtering them out improves relevance and search performance.
💡 Proposed Enhancement:
Thanks in advance for your feedback, and for all the great work on this project! 🙏