public-inbox.git - an "archives first" approach to mailing lists

Date	Commit message (Collapse)
2018-01-29	reply: follow obfuscation rules for HTML in sh args
	Namely, we do not want to obfuscate the mail address of the site itself.
2018-01-29	view: adjust wording for reply-to-list configs
	This makes the wording less confusing when showing archives for lists where the convention is reply-to-list. I still hate reply-to-list, but it's still better than no archives or list at all.
2018-01-26	atom: show metadata before message body
	This can allow streaming parsers (SAX) to work a little more efficiently as they can handle/discard all the metadata before the big content.
2018-01-16	hval: only allow domain obfuscation in address
	Obfuscating username portions of the email address leads to having subsequent parts of the address not being obfuscated; which could mean we show someone else's email entirely. In other words, obfuscating "john.doe@example.com" becomes might mean "doe@example.com" is picked up by scanners. In other news, email address obfuscation is still a horrible usability issue and only exists to appease misguided people.
2017-12-21	view: avoid deduping a single word in subject skeletons
	It is usually pointless to replace a single word with a '"' character.
2017-12-08	search: force large mbox result downloads to POST
	This should prevent crawlers (including most robots.txt ignoring ones) from burning our CPU time without severely compromising usability for humans.
2017-12-07	searchview: nofollow on mbox downloads
	Some search results are gigantic, and search engines are unlikely to be able to handle gzipped mboxes anyways.
2017-12-01	search: allow downloading search results as mbox
	Allowing downloading of all search results as an gzipped mboxrd file can be convenient for some users.
2017-11-29	view: avoid warning from negative repeat counts
	Perl 5.22 started warning about this.
2017-11-29	searchview: s/threaded/nested/
	We want to be consistent with the view change in commit b223e6f49debb99b9132bc85d97a065ebcee00b9
2017-11-16	watch: use "spam" in commit message for removals
	This makes it easy to identify the reason for message removals.
2017-11-16	learn: use "spam" as subject for removal commits
	Sometimes an email is an innocent removal "rm" for a misdirected, off-topic post, while most removed messages are "spam". Allow anybody to look at history and easily distinguish the reason for removing the message.
2017-10-18	view: s/threaded/nested/ in view
	We always do threading, so perhaps it's not a good name. "Nested" is probably more appropriate and closer to what people are used to seeing.
2017-10-04	mbox: support inline filename via Content-Disposition header
	This is hopefully more sensical than "raw" files from resulting downloads.
2017-10-03	search: try to fill in ghosts when generating thread skeleton
	Since we attempt to fill in threads by Subject, our thread skeletons can cross actual thread IDs, leading to the possibility of false ghosts showing up in the skeleton. Try to fill in the ghosts as well as possible by performing a message lookup.
2017-10-03	threading: deal with improperly-terminated References headers
	We should not blindly join References and In-Reply-To headers as a single string, because some messages can have an open angle brace '<' in References: without a corresponding '>'.
2017-07-13	www: Atom stream respects timezone
	Oops, we must not discard the timezone when parsing dates for the Atom stream.
2017-06-29	view: cull redundant phrases in subjects
	There is no need to show the same phrases over and over again in thread skeletons, it adds to visual noise and makes things more difficult to read.
2017-06-29	hval: only perform one substitution when obfuscating
	Only one substitution character is necessary when obfuscating email addresses.
2017-06-26	msgmap: reduce constant usage
	It is needless bloat and doesn't seem to help with readability, in retrospect, either.
2017-06-26	watch: avoid potential race condition while quitting
	We must not trigger future activity when initializing a -watch shutdown.
2017-06-26	watch: commit changes to fast-import sooner
	We should make changes visible sooner, even during lengthy scans.
2017-06-26	watch: use "self-inotify-tempfile trick" for quit
	This should be more reliable and safer as it'll ensure existing fast-import instances are shut down properly.
2017-06-26	watch: improve fairness during full rescans
	We need to ensure new messages are being processed fairly during full rescans, so have the ->scan subroutine yield and reschedule itself. Additionally, having a long-running task inside the signal handler is dangerous and subject to reentrancy bugs. Due to the limitations of the Filesys::Notify::Simple interface, we cannot rely on multiplexing I/O interfaces (select, IO::Poll, Danga::Socket, etc...) for this. Forking a separate process was considered, but it is more expensive for a mostly-idle process. So, we use a variant of the "self-pipe trick" via inotify (or whatever Filesys::Notify::Simple gives us). Instead of writing to our own pipe, we write to a file in our own temporary directory watched by Filesys::Notify::Simple to trigger events in signal handlers.
2017-06-26	spamc: retry on EINTR
	Signals can fire on us at any time if we're using blocking sysread.
2017-06-26	watch: ensure HUP causes the scanner to be reloaded
	Otherwise the old watcher may run indefinitely
2017-06-26	mda: set List-ID correctly according to RFC2919
	Oops, due to an old mistake , List-ID was set incorrectly in the MDA. This could cause some breakage w.r.t. mail filters.
2017-06-23	linkify: handle URLs in parenthesized statements
	Sometimes, URLs exist at the end of parethesized statements, and we shouldn't unnecessarily capture that. (example: https://public-inbox.org/ruby-core/20170623032722.GA8124@dcvr/)
2017-06-23	allow admins to configure non-obfuscated addresses/domains
	We will also treat all known list addresses as non-obfuscated. By setting publicinbox.noObfuscate in ~/.public-inbox/config, this will allow users to disable address obfuscation on a per-domain or per-address basis.
2017-06-23	config: assume lists have multiple addresses
	This should simplify the rest of our code for handling the do-not-obfuscate list.
2017-06-23	view: add newline before mailto: instructions in reply
	This is necessary to retain consistent spacing around bullet points. Fixes: 666844ae42b5b17f ("reply: handle address obfuscation :<")
2017-06-23	mbox: show application/mbox for obfuscated inboxes
	Sigh, yet another place to handle obfuscation for misguided people who expect it. Maybe this will do something to prevent spammers from getting addresses, while still allowing the "curl $URL \| git am" use case to work.
2017-06-23	reply: handle address obfuscation :<
	We can show users a lightly-obfuscated Bourne shell command for invoking "git send-email" for address obfuscation. However, I'm not sure if the mailto: arg will work effectively since URL encoding is probably too well-known to be effective.
2017-06-23	searchidx: fallback to lookup on pre-set article numbers
	Yet another hiccup from reusing pre-set article numbers on various ruby-lang.org mailing lists. This was causing messages to not appear to NNTP readers which use XOVER.
2017-06-23	msgmap: ignore duplicates instead of dying
	This prevents public-inbox-watch from dying when reloading (and thus rescanning) already-imported directories.
2017-06-23	watchmaildir: deal with rejected (100) messages
	The RubyLang filter is strict about what messages it rejects, so the spam learning path will not auto-train or remove messages missing X-Mail-Count headers.
2017-06-22	filter/rubylang: reuse altid entry from inbox object
	This allows users to DRY up their config a bit and avoid specifying altid twice when reusing the NNTP-centric msgmap for [ruby-*:\d+] serial numbers. My current work-in-progress ~/.public-inbox/config entry for the ruby-core list is: ------8<------- [publicinbox "ruby-core"] address = ruby-core@ruby-lang.org url = //public-inbox.org/ruby-core mainrepo = /path/to/ruby-core.git newsgroup = inbox.comp.lang.ruby.core watchheader = List-Id:<ruby-core.ruby-lang.org> altid = serial:ruby-core:file=msgmap.sqlite3 watch = maildir:/path/to/Maildir/.INBOX.ruby filter = PublicInbox::Filter::RubyLang
2017-06-22	msgmap: mid_insert ignores duplicates instead of die-ing
	This will allow smoother imports as occasional Message-ID duplicates happen and the best we can do is ignore the second one.
2017-06-22	add filter for RubyLang lists
	Unfortunately, it appears we have to reject this and instead add support filtering at View time(), due to DKIM signatures in messages from ruby-lang.org. () which may not be worth it
2017-06-20	import: fix encoding issues from weird "raw" emails
	This seems to allow weirdly-encoded "raw" emails in blade.nagaokaut.ac.jp/ruby/ruby-core/* to be handled without difficulties.
2017-06-16	view: implement optional address obfuscation
	This is lightly-tested and seems to work. I'm still hesitant to support this, but the alternative of receiving death threats for displaying unobfuscated addresses seems to be not worth it.
2017-06-15	reply: support Reply-To
	Reply-To is common and probably should've been supported, since day one, but we won't omit other addresses, either.
2017-06-15	replyto parameter support
	This allows us to support centralized mailing lists (which suck, but better than no mailing list at all).
2017-06-15	view: split out reply logic into its own module
	We'll be adding more reply options for centralized mailing lists. So split out the logic so it's easy-to-find. Organizing code is hard :<
2017-06-15	searchidx: remove messages correctly from Xapian index
	This fixes a bug introduced in commit 7eeadcb62729b0efbcb53cd9b7b181897c92cf9a ("search: remove unnecessary abstractions and functionality")
2017-06-14	search: allow searching within mail diffs
	This can be tied into a repository browser to browse in-flight topics on a mailing list.
2017-06-14	searchidx: switch to accounting by message bytes
	Xapian memory usage is tied to the size of the indexed text, so take the raw message size into account when deciding when to flush Xapian data. More importantly, we now flush Xapian before we have it buffer beyond our maximum; and we do it unconditionally to prevent even high priority processes from OOM-ing.
2017-06-14	search: remove unnecessary abstractions and functionality
	This simplifies the code a bit and reduces the translation overhead for looking directly at data from tools shipped with Xapian. While we're at it, fix thread-all.t :)
2017-06-07	filter/subjecttag: account for missing Subject: header
	This is a high indicator of spam (but out-of-scope for this particular module) but sometimes it is not, and people legitimately forget to set a Subject: header at all.
2017-05-25	import: reset :raw mode for commit title (subject)
	This was necessary for the presence of the 0xa0 byte() in the Subject: of the message at: http://blade.nagaokaut.ac.jp/ruby/ruby-core/3220 () That is 0xa0, not 0x0a ("\n"), so I wonder if the nibbles got swapped somehow.