You can not select more than 25 topics Topics must start with a letter or number, can include dashes ('-') and can be up to 35 characters long.

816 lines
22 KiB

Release 1.18 Changes are: * Update version number to 1.18 * Replace the basic fprintf call with a call to fwrite in order to work around the apparent compiler optimization/rewrite failure that we are seeing with the new toolchain/iOS SDKs provided with Xcode6 and iOS8. * Fix ALL the header guards. * Createed a README.md with the LevelDB project description. * A new CONTRIBUTING file. * Don't implicitly convert uint64_t to size_t or int. Either preserve it as uint64_t, or explicitly cast. This fixes MSVC warnings about possible value truncation when compiling this code in Chromium. * Added a DumpFile() library function that encapsulates the guts of the "leveldbutil dump" command. This will allow clients to dump data to their log files instead of stdout. It will also allow clients to supply their own environment. * leveldb: Remove unused function 'ConsumeChar'. * leveldbutil: Remove unused member variables from WriteBatchItemPrinter. * OpenBSD, NetBSD and DragonflyBSD have _LITTLE_ENDIAN, so define PLATFORM_IS_LITTLE_ENDIAN like on FreeBSD. This fixes: * issue #143 * issue #198 * issue #249 * Switch from <cstdatomic> to <atomic>. The former never made it into the standard and doesn't exist in modern gcc versions at all. The later contains everything that leveldb was using from the former. This problem was noticed when porting to Portable Native Client where no memory barrier is defined. The fact that <cstdatomic> is missing normally goes unnoticed since memory barriers are defined for most architectures. * Make Hash() treat its input as unsigned. Before this change LevelDB files from platforms with different signedness of char were not compatible. This change fixes: issue #243 * Verify checksums of index/meta/filter blocks when paranoid_checks set. * Invoke all tools for iOS with xcrun. (This was causing problems with the new XCode 5.1.1 image on pulse.) * include <sys/stat.h> only once, and fix the following linter warning: "Found C system header after C++ system header" * When encountering a corrupted table file, return Status::Corruption instead of Status::InvalidArgument. * Support cygwin as build platform, patch is from https://code.google.com/p/leveldb/issues/detail?id=188 * Fix typo, merge patch from https://code.google.com/p/leveldb/issues/detail?id=159 * Fix typos and comments, and address the following two issues: * issue #166 * issue #241 * Add missing db synchronize after "fillseq" in the benchmark. * Removed unused variable in SeekRandom: value (issue #201)
10 years ago
Release 1.18 Changes are: * Update version number to 1.18 * Replace the basic fprintf call with a call to fwrite in order to work around the apparent compiler optimization/rewrite failure that we are seeing with the new toolchain/iOS SDKs provided with Xcode6 and iOS8. * Fix ALL the header guards. * Createed a README.md with the LevelDB project description. * A new CONTRIBUTING file. * Don't implicitly convert uint64_t to size_t or int. Either preserve it as uint64_t, or explicitly cast. This fixes MSVC warnings about possible value truncation when compiling this code in Chromium. * Added a DumpFile() library function that encapsulates the guts of the "leveldbutil dump" command. This will allow clients to dump data to their log files instead of stdout. It will also allow clients to supply their own environment. * leveldb: Remove unused function 'ConsumeChar'. * leveldbutil: Remove unused member variables from WriteBatchItemPrinter. * OpenBSD, NetBSD and DragonflyBSD have _LITTLE_ENDIAN, so define PLATFORM_IS_LITTLE_ENDIAN like on FreeBSD. This fixes: * issue #143 * issue #198 * issue #249 * Switch from <cstdatomic> to <atomic>. The former never made it into the standard and doesn't exist in modern gcc versions at all. The later contains everything that leveldb was using from the former. This problem was noticed when porting to Portable Native Client where no memory barrier is defined. The fact that <cstdatomic> is missing normally goes unnoticed since memory barriers are defined for most architectures. * Make Hash() treat its input as unsigned. Before this change LevelDB files from platforms with different signedness of char were not compatible. This change fixes: issue #243 * Verify checksums of index/meta/filter blocks when paranoid_checks set. * Invoke all tools for iOS with xcrun. (This was causing problems with the new XCode 5.1.1 image on pulse.) * include <sys/stat.h> only once, and fix the following linter warning: "Found C system header after C++ system header" * When encountering a corrupted table file, return Status::Corruption instead of Status::InvalidArgument. * Support cygwin as build platform, patch is from https://code.google.com/p/leveldb/issues/detail?id=188 * Fix typo, merge patch from https://code.google.com/p/leveldb/issues/detail?id=159 * Fix typos and comments, and address the following two issues: * issue #166 * issue #241 * Add missing db synchronize after "fillseq" in the benchmark. * Removed unused variable in SeekRandom: value (issue #201)
10 years ago
  1. // Copyright (c) 2011 The LevelDB Authors. All rights reserved.
  2. // Use of this source code is governed by a BSD-style license that can be
  3. // found in the LICENSE file. See the AUTHORS file for names of contributors.
  4. #include <dirent.h>
  5. #include <errno.h>
  6. #include <fcntl.h>
  7. #include <pthread.h>
  8. #include <stdlib.h>
  9. #include <string.h>
  10. #include <sys/mman.h>
  11. #include <sys/resource.h>
  12. #include <sys/stat.h>
  13. #include <sys/time.h>
  14. #include <sys/types.h>
  15. #include <time.h>
  16. #include <unistd.h>
  17. #include <atomic>
  18. #include <cstddef>
  19. #include <cstdint>
  20. #include <cstring>
  21. #include <limits>
  22. #include <queue>
  23. #include <set>
  24. #include <string>
  25. #include <thread>
  26. #include <type_traits>
  27. #include "leveldb/env.h"
  28. #include "leveldb/slice.h"
  29. #include "leveldb/status.h"
  30. #include "port/port.h"
  31. #include "port/thread_annotations.h"
  32. #include "util/posix_logger.h"
  33. #include "util/env_posix_test_helper.h"
  34. // HAVE_FDATASYNC is defined in the auto-generated port_config.h, which is
  35. // included by port_stdcxx.h.
  36. #if !HAVE_FDATASYNC
  37. #define fdatasync fsync
  38. #endif // !HAVE_FDATASYNC
  39. namespace leveldb {
  40. namespace {
  41. static int open_read_only_file_limit = -1;
  42. static int mmap_limit = -1;
  43. constexpr const size_t kWritableFileBufferSize = 65536;
  44. static Status PosixError(const std::string& context, int err_number) {
  45. if (err_number == ENOENT) {
  46. return Status::NotFound(context, strerror(err_number));
  47. } else {
  48. return Status::IOError(context, strerror(err_number));
  49. }
  50. }
  51. // Helper class to limit resource usage to avoid exhaustion.
  52. // Currently used to limit read-only file descriptors and mmap file usage
  53. // so that we do not run out of file descriptors or virtual memory, or run into
  54. // kernel performance problems for very large databases.
  55. class Limiter {
  56. public:
  57. // Limit maximum number of resources to |max_acquires|.
  58. Limiter(int max_acquires) : acquires_allowed_(max_acquires) {}
  59. Limiter(const Limiter&) = delete;
  60. Limiter operator=(const Limiter&) = delete;
  61. // If another resource is available, acquire it and return true.
  62. // Else return false.
  63. bool Acquire() {
  64. int old_acquires_allowed =
  65. acquires_allowed_.fetch_sub(1, std::memory_order_relaxed);
  66. if (old_acquires_allowed > 0)
  67. return true;
  68. acquires_allowed_.fetch_add(1, std::memory_order_relaxed);
  69. return false;
  70. }
  71. // Release a resource acquired by a previous call to Acquire() that returned
  72. // true.
  73. void Release() {
  74. acquires_allowed_.fetch_add(1, std::memory_order_relaxed);
  75. }
  76. private:
  77. // The number of available resources.
  78. //
  79. // This is a counter and is not tied to the invariants of any other class, so
  80. // it can be operated on safely using std::memory_order_relaxed.
  81. std::atomic<int> acquires_allowed_;
  82. };
  83. class PosixSequentialFile: public SequentialFile {
  84. private:
  85. std::string filename_;
  86. int fd_;
  87. public:
  88. PosixSequentialFile(const std::string& fname, int fd)
  89. : filename_(fname), fd_(fd) {}
  90. virtual ~PosixSequentialFile() { close(fd_); }
  91. virtual Status Read(size_t n, Slice* result, char* scratch) {
  92. Status s;
  93. while (true) {
  94. ssize_t r = read(fd_, scratch, n);
  95. if (r < 0) {
  96. if (errno == EINTR) {
  97. continue; // Retry
  98. }
  99. s = PosixError(filename_, errno);
  100. break;
  101. }
  102. *result = Slice(scratch, r);
  103. break;
  104. }
  105. return s;
  106. }
  107. virtual Status Skip(uint64_t n) {
  108. if (lseek(fd_, n, SEEK_CUR) == static_cast<off_t>(-1)) {
  109. return PosixError(filename_, errno);
  110. }
  111. return Status::OK();
  112. }
  113. };
  114. // pread() based random-access
  115. class PosixRandomAccessFile: public RandomAccessFile {
  116. private:
  117. std::string filename_;
  118. bool temporary_fd_; // If true, fd_ is -1 and we open on every read.
  119. int fd_;
  120. Limiter* limiter_;
  121. public:
  122. PosixRandomAccessFile(const std::string& fname, int fd, Limiter* limiter)
  123. : filename_(fname), fd_(fd), limiter_(limiter) {
  124. temporary_fd_ = !limiter->Acquire();
  125. if (temporary_fd_) {
  126. // Open file on every access.
  127. close(fd_);
  128. fd_ = -1;
  129. }
  130. }
  131. virtual ~PosixRandomAccessFile() {
  132. if (!temporary_fd_) {
  133. close(fd_);
  134. limiter_->Release();
  135. }
  136. }
  137. virtual Status Read(uint64_t offset, size_t n, Slice* result,
  138. char* scratch) const {
  139. int fd = fd_;
  140. if (temporary_fd_) {
  141. fd = open(filename_.c_str(), O_RDONLY);
  142. if (fd < 0) {
  143. return PosixError(filename_, errno);
  144. }
  145. }
  146. Status s;
  147. ssize_t r = pread(fd, scratch, n, static_cast<off_t>(offset));
  148. *result = Slice(scratch, (r < 0) ? 0 : r);
  149. if (r < 0) {
  150. // An error: return a non-ok status
  151. s = PosixError(filename_, errno);
  152. }
  153. if (temporary_fd_) {
  154. // Close the temporary file descriptor opened earlier.
  155. close(fd);
  156. }
  157. return s;
  158. }
  159. };
  160. // mmap() based random-access
  161. class PosixMmapReadableFile: public RandomAccessFile {
  162. private:
  163. std::string filename_;
  164. void* mmapped_region_;
  165. size_t length_;
  166. Limiter* limiter_;
  167. public:
  168. // base[0,length-1] contains the mmapped contents of the file.
  169. PosixMmapReadableFile(const std::string& fname, void* base, size_t length,
  170. Limiter* limiter)
  171. : filename_(fname), mmapped_region_(base), length_(length),
  172. limiter_(limiter) {
  173. }
  174. virtual ~PosixMmapReadableFile() {
  175. munmap(mmapped_region_, length_);
  176. limiter_->Release();
  177. }
  178. virtual Status Read(uint64_t offset, size_t n, Slice* result,
  179. char* scratch) const {
  180. Status s;
  181. if (offset + n > length_) {
  182. *result = Slice();
  183. s = PosixError(filename_, EINVAL);
  184. } else {
  185. *result = Slice(reinterpret_cast<char*>(mmapped_region_) + offset, n);
  186. }
  187. return s;
  188. }
  189. };
  190. class PosixWritableFile final : public WritableFile {
  191. public:
  192. PosixWritableFile(std::string filename, int fd)
  193. : pos_(0), fd_(fd), is_manifest_(IsManifest(filename)),
  194. filename_(std::move(filename)), dirname_(Dirname(filename_)) {}
  195. ~PosixWritableFile() override {
  196. if (fd_ >= 0) {
  197. // Ignoring any potential errors
  198. Close();
  199. }
  200. }
  201. Status Append(const Slice& data) override {
  202. size_t write_size = data.size();
  203. const char* write_data = data.data();
  204. // Fit as much as possible into buffer.
  205. size_t copy_size = std::min(write_size, kWritableFileBufferSize - pos_);
  206. std::memcpy(buf_ + pos_, write_data, copy_size);
  207. write_data += copy_size;
  208. write_size -= copy_size;
  209. pos_ += copy_size;
  210. if (write_size == 0) {
  211. return Status::OK();
  212. }
  213. // Can't fit in buffer, so need to do at least one write.
  214. Status status = FlushBuffer();
  215. if (!status.ok()) {
  216. return status;
  217. }
  218. // Small writes go to buffer, large writes are written directly.
  219. if (write_size < kWritableFileBufferSize) {
  220. std::memcpy(buf_, write_data, write_size);
  221. pos_ = write_size;
  222. return Status::OK();
  223. }
  224. return WriteUnbuffered(write_data, write_size);
  225. }
  226. Status Close() override {
  227. Status status = FlushBuffer();
  228. const int close_result = ::close(fd_);
  229. if (close_result < 0 && status.ok()) {
  230. status = PosixError(filename_, errno);
  231. }
  232. fd_ = -1;
  233. return status;
  234. }
  235. Status Flush() override {
  236. return FlushBuffer();
  237. }
  238. Status Sync() override {
  239. // Ensure new files referred to by the manifest are in the filesystem.
  240. //
  241. // This needs to happen before the manifest file is flushed to disk, to
  242. // avoid crashing in a state where the manifest refers to files that are not
  243. // yet on disk.
  244. Status status = SyncDirIfManifest();
  245. if (!status.ok()) {
  246. return status;
  247. }
  248. status = FlushBuffer();
  249. if (status.ok() && ::fdatasync(fd_) != 0) {
  250. status = PosixError(filename_, errno);
  251. }
  252. return status;
  253. }
  254. private:
  255. Status FlushBuffer() {
  256. Status status = WriteUnbuffered(buf_, pos_);
  257. pos_ = 0;
  258. return status;
  259. }
  260. Status WriteUnbuffered(const char* data, size_t size) {
  261. while (size > 0) {
  262. ssize_t write_result = ::write(fd_, data, size);
  263. if (write_result < 0) {
  264. if (errno == EINTR) {
  265. continue; // Retry
  266. }
  267. return PosixError(filename_, errno);
  268. }
  269. data += write_result;
  270. size -= write_result;
  271. }
  272. return Status::OK();
  273. }
  274. Status SyncDirIfManifest() {
  275. Status status;
  276. if (!is_manifest_) {
  277. return status;
  278. }
  279. int fd = ::open(dirname_.c_str(), O_RDONLY);
  280. if (fd < 0) {
  281. status = PosixError(dirname_, errno);
  282. } else {
  283. if (::fsync(fd) < 0) {
  284. status = PosixError(dirname_, errno);
  285. }
  286. ::close(fd);
  287. }
  288. return status;
  289. }
  290. // Returns the directory name in a path pointing to a file.
  291. //
  292. // Returns "." if the path does not contain any directory separator.
  293. static std::string Dirname(const std::string& filename) {
  294. std::string::size_type separator_pos = filename.rfind('/');
  295. if (separator_pos == std::string::npos) {
  296. return std::string(".");
  297. }
  298. // The filename component should not contain a path separator. If it does,
  299. // the splitting was done incorrectly.
  300. assert(filename.find('/', separator_pos + 1) == std::string::npos);
  301. return filename.substr(0, separator_pos);
  302. }
  303. // Extracts the file name from a path pointing to a file.
  304. //
  305. // The returned Slice points to |filename|'s data buffer, so it is only valid
  306. // while |filename| is alive and unchanged.
  307. static Slice Basename(const std::string& filename) {
  308. std::string::size_type separator_pos = filename.rfind('/');
  309. if (separator_pos == std::string::npos) {
  310. return Slice(filename);
  311. }
  312. // The filename component should not contain a path separator. If it does,
  313. // the splitting was done incorrectly.
  314. assert(filename.find('/', separator_pos + 1) == std::string::npos);
  315. return Slice(filename.data() + separator_pos + 1,
  316. filename.length() - separator_pos - 1);
  317. }
  318. // True if the given file is a manifest file.
  319. static bool IsManifest(const std::string& filename) {
  320. return Basename(filename).starts_with("MANIFEST");
  321. }
  322. // buf_[0, pos_ - 1] contains data to be written to fd_.
  323. char buf_[kWritableFileBufferSize];
  324. size_t pos_;
  325. int fd_;
  326. const bool is_manifest_; // True if the file's name starts with MANIFEST.
  327. const std::string filename_;
  328. const std::string dirname_; // The directory of filename_.
  329. };
  330. static int LockOrUnlock(int fd, bool lock) {
  331. errno = 0;
  332. struct flock f;
  333. memset(&f, 0, sizeof(f));
  334. f.l_type = (lock ? F_WRLCK : F_UNLCK);
  335. f.l_whence = SEEK_SET;
  336. f.l_start = 0;
  337. f.l_len = 0; // Lock/unlock entire file
  338. return fcntl(fd, F_SETLK, &f);
  339. }
  340. class PosixFileLock : public FileLock {
  341. public:
  342. int fd_;
  343. std::string name_;
  344. };
  345. // Set of locked files. We keep a separate set instead of just
  346. // relying on fcntrl(F_SETLK) since fcntl(F_SETLK) does not provide
  347. // any protection against multiple uses from the same process.
  348. class PosixLockTable {
  349. private:
  350. port::Mutex mu_;
  351. std::set<std::string> locked_files_ GUARDED_BY(mu_);
  352. public:
  353. bool Insert(const std::string& fname) LOCKS_EXCLUDED(mu_) {
  354. mu_.Lock();
  355. bool succeeded = locked_files_.insert(fname).second;
  356. mu_.Unlock();
  357. return succeeded;
  358. }
  359. void Remove(const std::string& fname) LOCKS_EXCLUDED(mu_) {
  360. mu_.Lock();
  361. locked_files_.erase(fname);
  362. mu_.Unlock();
  363. }
  364. };
  365. class PosixEnv : public Env {
  366. public:
  367. PosixEnv();
  368. virtual ~PosixEnv() {
  369. char msg[] = "Destroying Env::Default()\n";
  370. fwrite(msg, 1, sizeof(msg), stderr);
  371. abort();
  372. }
  373. virtual Status NewSequentialFile(const std::string& fname,
  374. SequentialFile** result) {
  375. int fd = open(fname.c_str(), O_RDONLY);
  376. if (fd < 0) {
  377. *result = nullptr;
  378. return PosixError(fname, errno);
  379. } else {
  380. *result = new PosixSequentialFile(fname, fd);
  381. return Status::OK();
  382. }
  383. }
  384. virtual Status NewRandomAccessFile(const std::string& fname,
  385. RandomAccessFile** result) {
  386. *result = nullptr;
  387. Status s;
  388. int fd = open(fname.c_str(), O_RDONLY);
  389. if (fd < 0) {
  390. s = PosixError(fname, errno);
  391. } else if (mmap_limit_.Acquire()) {
  392. uint64_t size;
  393. s = GetFileSize(fname, &size);
  394. if (s.ok()) {
  395. void* base = mmap(nullptr, size, PROT_READ, MAP_SHARED, fd, 0);
  396. if (base != MAP_FAILED) {
  397. *result = new PosixMmapReadableFile(fname, base, size, &mmap_limit_);
  398. } else {
  399. s = PosixError(fname, errno);
  400. }
  401. }
  402. close(fd);
  403. if (!s.ok()) {
  404. mmap_limit_.Release();
  405. }
  406. } else {
  407. *result = new PosixRandomAccessFile(fname, fd, &fd_limit_);
  408. }
  409. return s;
  410. }
  411. virtual Status NewWritableFile(const std::string& fname,
  412. WritableFile** result) {
  413. Status s;
  414. int fd = open(fname.c_str(), O_TRUNC | O_WRONLY | O_CREAT, 0644);
  415. if (fd < 0) {
  416. *result = nullptr;
  417. s = PosixError(fname, errno);
  418. } else {
  419. *result = new PosixWritableFile(fname, fd);
  420. }
  421. return s;
  422. }
  423. virtual Status NewAppendableFile(const std::string& fname,
  424. WritableFile** result) {
  425. Status s;
  426. int fd = open(fname.c_str(), O_APPEND | O_WRONLY | O_CREAT, 0644);
  427. if (fd < 0) {
  428. *result = nullptr;
  429. s = PosixError(fname, errno);
  430. } else {
  431. *result = new PosixWritableFile(fname, fd);
  432. }
  433. return s;
  434. }
  435. virtual bool FileExists(const std::string& fname) {
  436. return access(fname.c_str(), F_OK) == 0;
  437. }
  438. virtual Status GetChildren(const std::string& dir,
  439. std::vector<std::string>* result) {
  440. result->clear();
  441. DIR* d = opendir(dir.c_str());
  442. if (d == nullptr) {
  443. return PosixError(dir, errno);
  444. }
  445. struct dirent* entry;
  446. while ((entry = readdir(d)) != nullptr) {
  447. result->push_back(entry->d_name);
  448. }
  449. closedir(d);
  450. return Status::OK();
  451. }
  452. virtual Status DeleteFile(const std::string& fname) {
  453. Status result;
  454. if (unlink(fname.c_str()) != 0) {
  455. result = PosixError(fname, errno);
  456. }
  457. return result;
  458. }
  459. virtual Status CreateDir(const std::string& name) {
  460. Status result;
  461. if (mkdir(name.c_str(), 0755) != 0) {
  462. result = PosixError(name, errno);
  463. }
  464. return result;
  465. }
  466. virtual Status DeleteDir(const std::string& name) {
  467. Status result;
  468. if (rmdir(name.c_str()) != 0) {
  469. result = PosixError(name, errno);
  470. }
  471. return result;
  472. }
  473. virtual Status GetFileSize(const std::string& fname, uint64_t* size) {
  474. Status s;
  475. struct stat sbuf;
  476. if (stat(fname.c_str(), &sbuf) != 0) {
  477. *size = 0;
  478. s = PosixError(fname, errno);
  479. } else {
  480. *size = sbuf.st_size;
  481. }
  482. return s;
  483. }
  484. virtual Status RenameFile(const std::string& src, const std::string& target) {
  485. Status result;
  486. if (rename(src.c_str(), target.c_str()) != 0) {
  487. result = PosixError(src, errno);
  488. }
  489. return result;
  490. }
  491. virtual Status LockFile(const std::string& fname, FileLock** lock) {
  492. *lock = nullptr;
  493. Status result;
  494. int fd = open(fname.c_str(), O_RDWR | O_CREAT, 0644);
  495. if (fd < 0) {
  496. result = PosixError(fname, errno);
  497. } else if (!locks_.Insert(fname)) {
  498. close(fd);
  499. result = Status::IOError("lock " + fname, "already held by process");
  500. } else if (LockOrUnlock(fd, true) == -1) {
  501. result = PosixError("lock " + fname, errno);
  502. close(fd);
  503. locks_.Remove(fname);
  504. } else {
  505. PosixFileLock* my_lock = new PosixFileLock;
  506. my_lock->fd_ = fd;
  507. my_lock->name_ = fname;
  508. *lock = my_lock;
  509. }
  510. return result;
  511. }
  512. virtual Status UnlockFile(FileLock* lock) {
  513. PosixFileLock* my_lock = reinterpret_cast<PosixFileLock*>(lock);
  514. Status result;
  515. if (LockOrUnlock(my_lock->fd_, false) == -1) {
  516. result = PosixError("unlock", errno);
  517. }
  518. locks_.Remove(my_lock->name_);
  519. close(my_lock->fd_);
  520. delete my_lock;
  521. return result;
  522. }
  523. virtual void Schedule(void (*function)(void*), void* arg);
  524. virtual void StartThread(void (*function)(void* arg), void* arg);
  525. virtual Status GetTestDirectory(std::string* result) {
  526. const char* env = getenv("TEST_TMPDIR");
  527. if (env && env[0] != '\0') {
  528. *result = env;
  529. } else {
  530. char buf[100];
  531. snprintf(buf, sizeof(buf), "/tmp/leveldbtest-%d", int(geteuid()));
  532. *result = buf;
  533. }
  534. // Directory may already exist
  535. CreateDir(*result);
  536. return Status::OK();
  537. }
  538. virtual Status NewLogger(const std::string& fname, Logger** result) {
  539. FILE* f = fopen(fname.c_str(), "w");
  540. if (f == nullptr) {
  541. *result = nullptr;
  542. return PosixError(fname, errno);
  543. } else {
  544. *result = new PosixLogger(f);
  545. return Status::OK();
  546. }
  547. }
  548. virtual uint64_t NowMicros() {
  549. struct timeval tv;
  550. gettimeofday(&tv, nullptr);
  551. return static_cast<uint64_t>(tv.tv_sec) * 1000000 + tv.tv_usec;
  552. }
  553. virtual void SleepForMicroseconds(int micros) {
  554. usleep(micros);
  555. }
  556. private:
  557. void BackgroundThreadMain();
  558. static void BackgroundThreadEntryPoint(PosixEnv* env) {
  559. env->BackgroundThreadMain();
  560. }
  561. // Stores the work item data in a Schedule() call.
  562. //
  563. // Instances are constructed on the thread calling Schedule() and used on the
  564. // background thread.
  565. //
  566. // This structure is thread-safe beacuse it is immutable.
  567. struct BackgroundWorkItem {
  568. explicit BackgroundWorkItem(void (*function)(void* arg), void* arg)
  569. : function(function), arg(arg) {}
  570. void (* const function)(void*);
  571. void* const arg;
  572. };
  573. port::Mutex background_work_mutex_;
  574. port::CondVar background_work_cv_ GUARDED_BY(background_work_mutex_);
  575. bool started_background_thread_ GUARDED_BY(background_work_mutex_);
  576. std::queue<BackgroundWorkItem> background_work_queue_
  577. GUARDED_BY(background_work_mutex_);
  578. PosixLockTable locks_;
  579. Limiter mmap_limit_;
  580. Limiter fd_limit_;
  581. };
  582. // Return the maximum number of concurrent mmaps.
  583. static int MaxMmaps() {
  584. if (mmap_limit >= 0) {
  585. return mmap_limit;
  586. }
  587. // Up to 1000 mmaps for 64-bit binaries; none for smaller pointer sizes.
  588. mmap_limit = sizeof(void*) >= 8 ? 1000 : 0;
  589. return mmap_limit;
  590. }
  591. // Return the maximum number of read-only files to keep open.
  592. static intptr_t MaxOpenFiles() {
  593. if (open_read_only_file_limit >= 0) {
  594. return open_read_only_file_limit;
  595. }
  596. struct rlimit rlim;
  597. if (getrlimit(RLIMIT_NOFILE, &rlim)) {
  598. // getrlimit failed, fallback to hard-coded default.
  599. open_read_only_file_limit = 50;
  600. } else if (rlim.rlim_cur == RLIM_INFINITY) {
  601. open_read_only_file_limit = std::numeric_limits<int>::max();
  602. } else {
  603. // Allow use of 20% of available file descriptors for read-only files.
  604. open_read_only_file_limit = rlim.rlim_cur / 5;
  605. }
  606. return open_read_only_file_limit;
  607. }
  608. PosixEnv::PosixEnv()
  609. : background_work_cv_(&background_work_mutex_),
  610. started_background_thread_(false),
  611. mmap_limit_(MaxMmaps()),
  612. fd_limit_(MaxOpenFiles()) {
  613. }
  614. void PosixEnv::Schedule(
  615. void (*background_work_function)(void* background_work_arg),
  616. void* background_work_arg) {
  617. background_work_mutex_.Lock();
  618. // Start the background thread, if we haven't done so already.
  619. if (!started_background_thread_) {
  620. started_background_thread_ = true;
  621. std::thread background_thread(PosixEnv::BackgroundThreadEntryPoint, this);
  622. background_thread.detach();
  623. }
  624. // If the queue is empty, the background thread may be waiting for work.
  625. if (background_work_queue_.empty()) {
  626. background_work_cv_.Signal();
  627. }
  628. background_work_queue_.emplace(background_work_function, background_work_arg);
  629. background_work_mutex_.Unlock();
  630. }
  631. void PosixEnv::BackgroundThreadMain() {
  632. while (true) {
  633. background_work_mutex_.Lock();
  634. // Wait until there is work to be done.
  635. while (background_work_queue_.empty()) {
  636. background_work_cv_.Wait();
  637. }
  638. assert(!background_work_queue_.empty());
  639. auto background_work_function =
  640. background_work_queue_.front().function;
  641. void* background_work_arg = background_work_queue_.front().arg;
  642. background_work_queue_.pop();
  643. background_work_mutex_.Unlock();
  644. background_work_function(background_work_arg);
  645. }
  646. }
  647. // Wraps an Env instance whose destructor is never created.
  648. //
  649. // Intended usage:
  650. // using PlatformSingletonEnv = SingletonEnv<PlatformEnv>;
  651. // void ConfigurePosixEnv(int param) {
  652. // PlatformSingletonEnv::AssertEnvNotInitialized();
  653. // // set global configuration flags.
  654. // }
  655. // Env* Env::Default() {
  656. // static PlatformSingletonEnv default_env;
  657. // return default_env.env();
  658. // }
  659. template<typename EnvType>
  660. class SingletonEnv {
  661. public:
  662. SingletonEnv() {
  663. #if !defined(NDEBUG)
  664. env_initialized_.store(true, std::memory_order::memory_order_relaxed);
  665. #endif // !defined(NDEBUG)
  666. static_assert(sizeof(env_storage_) >= sizeof(EnvType),
  667. "env_storage_ will not fit the Env");
  668. static_assert(alignof(decltype(env_storage_)) >= alignof(EnvType),
  669. "env_storage_ does not meet the Env's alignment needs");
  670. new (&env_storage_) EnvType();
  671. }
  672. ~SingletonEnv() = default;
  673. SingletonEnv(const SingletonEnv&) = delete;
  674. SingletonEnv& operator=(const SingletonEnv&) = delete;
  675. Env* env() { return reinterpret_cast<Env*>(&env_storage_); }
  676. static void AssertEnvNotInitialized() {
  677. #if !defined(NDEBUG)
  678. assert(!env_initialized_.load(std::memory_order::memory_order_relaxed));
  679. #endif // !defined(NDEBUG)
  680. }
  681. private:
  682. typename std::aligned_storage<sizeof(EnvType), alignof(EnvType)>::type
  683. env_storage_;
  684. #if !defined(NDEBUG)
  685. static std::atomic<bool> env_initialized_;
  686. #endif // !defined(NDEBUG)
  687. };
  688. #if !defined(NDEBUG)
  689. template<typename EnvType>
  690. std::atomic<bool> SingletonEnv<EnvType>::env_initialized_;
  691. #endif // !defined(NDEBUG)
  692. using PosixDefaultEnv = SingletonEnv<PosixEnv>;
  693. } // namespace
  694. void PosixEnv::StartThread(void (*thread_main)(void* thread_main_arg),
  695. void* thread_main_arg) {
  696. std::thread new_thread(thread_main, thread_main_arg);
  697. new_thread.detach();
  698. }
  699. void EnvPosixTestHelper::SetReadOnlyFDLimit(int limit) {
  700. PosixDefaultEnv::AssertEnvNotInitialized();
  701. open_read_only_file_limit = limit;
  702. }
  703. void EnvPosixTestHelper::SetReadOnlyMMapLimit(int limit) {
  704. PosixDefaultEnv::AssertEnvNotInitialized();
  705. mmap_limit = limit;
  706. }
  707. Env* Env::Default() {
  708. static PosixDefaultEnv env_container;
  709. return env_container.env();
  710. }
  711. } // namespace leveldb