From: Eric Wong <e@yhbt.net>
To: meta@public-inbox.org
Subject: [PATCH 00/20] indexing changes and new features
Date: Fri, 24 Jul 2020 05:55:46 +0000 [thread overview]
Message-ID: <20200724055606.27332-1-e@yhbt.net> (raw)
--rethread and --no-sync options are now supported in
public-inbox-index. --no-sync should be nice for users
of FSes with poor fsync(2) performance.
Now: I also wonder if --no-sync is a bad name since we
also use it for to mean synchronising indices. Perhaps
--no-fsync would be a better name, though technically
SQLite and Xapian use fdatasync(2), nowadays.
Some of this is prep work for exposing THREADID via IMAP (and
JMAP) to aid in searching.
Since THREADID (`over.tid') will be exposed in a user-visible
way, I'm finally giving up on using the default (reverse
chronological) log order for indexing to ensure THREADID
ascends for newer threads.
This also simplifies the indexing code significantly.
To avoid pinning huge amounts of RAM, the working space is held
in a IdxStack temporary file. This further simplifies our code
since we no longer have to worry about old that did not use
Xapian w/o FD_CLOEXEC.
There's still more work on the horizon, here...
Eric Wong (20):
index: support --rethread switch to fix old indices
v2: index forwards (via `git log --reverse')
v2writable: introduce idx_stack
v2writable: index_sync: reduce fill_alternates calls
v2writable: move {autime} and {cotime} into $sync state
v2writable: allow >= 40 byte git object IDs
v2writable: drop "EPOCH.git indexing $RANGE" progress message
use consistent {ibx} field for writable code paths
search: avoid copying {inboxdir}
v2writable: use read-only PublicInbox::Git for cat_file
v2writable: get rid of {reindex_pipe} field
v2writable: clarify "epoch" for {last_commits}
xapcmd: set {from} properly for v1 inboxes
searchidx: rename _xdb_{acquire,release} => idx_
searchidx: make v1 indexing closer to v2
index+xcpdb: support --no-sync flag
v2writable: share log2stack code with v1
searchidx: support async git check
searchidx: $batch_cb => v1_checkpoint
v2writable: {unindexed} belongs in $sync state
Documentation/public-inbox-index.pod | 30 +-
Documentation/public-inbox-xcpdb.pod | 6 +
MANIFEST | 3 +-
lib/PublicInbox/Git.pm | 72 ++++-
lib/PublicInbox/IdxStack.pm | 52 ++++
lib/PublicInbox/Import.pm | 6 +-
lib/PublicInbox/Msgmap.pm | 21 +-
lib/PublicInbox/MultiMidQueue.pm | 62 ----
lib/PublicInbox/Over.pm | 1 +
lib/PublicInbox/OverIdx.pm | 78 ++++-
lib/PublicInbox/Search.pm | 25 +-
lib/PublicInbox/SearchIdx.pm | 384 ++++++++++++------------
lib/PublicInbox/SearchIdxShard.pm | 12 +-
lib/PublicInbox/Smsg.pm | 8 +-
lib/PublicInbox/V2Writable.pm | 427 +++++++++------------------
lib/PublicInbox/Xapcmd.pm | 10 +-
script/public-inbox-index | 5 +-
script/public-inbox-xcpdb | 4 +-
t/idx_stack.t | 56 ++++
t/inbox_idle.t | 4 +-
t/search.t | 4 +-
t/v1reindex.t | 36 ++-
t/v2reindex.t | 45 +++
23 files changed, 744 insertions(+), 607 deletions(-)
create mode 100644 lib/PublicInbox/IdxStack.pm
delete mode 100644 lib/PublicInbox/MultiMidQueue.pm
create mode 100644 t/idx_stack.t
next reply other threads:[~2020-07-24 5:56 UTC|newest]
Thread overview: 21+ messages / expand[flat|nested] mbox.gz Atom feed top
2020-07-24 5:55 Eric Wong [this message]
2020-07-24 5:55 ` [PATCH 01/20] index: support --rethread switch to fix old indices Eric Wong
2020-07-24 5:55 ` [PATCH 02/20] v2: index forwards (via `git log --reverse') Eric Wong
2020-07-24 5:55 ` [PATCH 03/20] v2writable: introduce idx_stack Eric Wong
2020-07-24 5:55 ` [PATCH 04/20] v2writable: index_sync: reduce fill_alternates calls Eric Wong
2020-07-24 5:55 ` [PATCH 05/20] v2writable: move {autime} and {cotime} into $sync state Eric Wong
2020-07-24 5:55 ` [PATCH 06/20] v2writable: allow >= 40 byte git object IDs Eric Wong
2020-07-24 5:55 ` [PATCH 07/20] v2writable: drop "EPOCH.git indexing $RANGE" progress Eric Wong
2020-07-24 5:55 ` [PATCH 08/20] use consistent {ibx} field for writable code paths Eric Wong
2020-07-24 5:55 ` [PATCH 09/20] search: avoid copying {inboxdir} Eric Wong
2020-07-24 5:55 ` [PATCH 10/20] v2writable: use read-only PublicInbox::Git for cat_file Eric Wong
2020-07-24 5:55 ` [PATCH 11/20] v2writable: get rid of {reindex_pipe} field Eric Wong
2020-07-24 5:55 ` [PATCH 12/20] v2writable: clarify "epoch" comment Eric Wong
2020-07-24 5:55 ` [PATCH 13/20] xapcmd: set {from} properly for v1 inboxes Eric Wong
2020-07-24 5:56 ` [PATCH 14/20] searchidx: rename _xdb_{acquire,release} => idx_ Eric Wong
2020-07-24 5:56 ` [PATCH 15/20] searchidx: make v1 indexing closer to v2 Eric Wong
2020-07-24 5:56 ` [PATCH 16/20] index+xcpdb: support --no-sync flag Eric Wong
2020-07-24 5:56 ` [PATCH 17/20] v2writable: share log2stack code with v1 Eric Wong
2020-07-24 5:56 ` [PATCH 18/20] searchidx: support async git check Eric Wong
2020-07-24 5:56 ` [PATCH 19/20] searchidx: $batch_cb => v1_checkpoint Eric Wong
2020-07-24 5:56 ` [PATCH 20/20] v2writable: {unindexed} belongs in $sync state Eric Wong
Reply instructions:
You may reply publicly to this message via plain-text email
using any one of the following methods:
* Save the following mbox file, import it into your mail client,
and reply-to-all from there: mbox
Avoid top-posting and favor interleaved quoting:
https://en.wikipedia.org/wiki/Posting_style#Interleaved_style
List information: https://public-inbox.org/README
* Reply using the --to, --cc, and --in-reply-to
switches of git-send-email(1):
git send-email \
--in-reply-to=20200724055606.27332-1-e@yhbt.net \
--to=e@yhbt.net \
--cc=meta@public-inbox.org \
/path/to/YOUR_REPLY
https://kernel.org/pub/software/scm/git/docs/git-send-email.html
* If your mail client supports setting the In-Reply-To header
via mailto: links, try the mailto: link
Be sure your reply has a Subject: header at the top and a blank line
before the message body.
Code repositories for project(s) associated with this public inbox
https://80x24.org/public-inbox.git
This is a public inbox, see mirroring instructions
for how to clone and mirror all data and code used for this inbox;
as well as URLs for read-only IMAP folder(s) and NNTP newsgroup(s).