diff --git a/RefactorPlan b/RefactorPlan index ef776353..44a24cfb 100644 --- a/RefactorPlan +++ b/RefactorPlan @@ -82,6 +82,40 @@ and repository name as key and add that dictionary to db. For example if I am ad packages db, I am actually adding, a Dictionary with ["repotype", foo] to db. This is cheap, bloated and slow. We must have different repository databases for packages. +* We should really consider using a simple and fast relational database (sqlite) for database operations. Here +is the rationale and some issues. +- Current implementation uses berkeleydb as a storage of dictionaries, with keys and Pickled Python objects. +This approach simplifies coding requirements but leads to some problems. +1. Some operations cannot be done effectively. Searching can only be done by iterating all objects and searching +one by one. component, reverse dependencies etc are holding lists of package names that are updated on every +package operation, Multiple repository problem is + +2.Performance: although fast, pickling and unpickling whole objects is still a time consuming process. +Pickled Objects contains object metadata and are larger then the data they hold, so reading, writing pickled +Objects to db is slower than reading and writing only necessary fields. large objects (Order of kilobytes) +causes berkeleydb data buckets to overflow and caches works less efficently. + +4.Disk usage: Using pickled objects requires more memory and disk space. Total size of a current pisi database +is over 100Mb, if we modify files.bdb to only hold filenames and packages, size is still around 50MB. + +Rational Databases offers number of advantages and disadvantages: + +Advantages: ++ Relational databases offers much more flexible operations on stored data. Searching, grouping, sorting +can be done by database engine efficently. ++ There will be only a single database file and much smaller database size compared to berkeleydb files +holding pickled python objects. ++ Since database no longer holds pickled python objects, it is possible to read and modify the database +by other programs written in different languages (Java, C++ etc.) ++ SQL lite is generally very fast, probably beating "Berkeleydb with pickled objects" by the order of 4-5 on +queries. + +Disadvantages +- Using a relational database will require more coding when adding and retrieving contents of objects. +- Introducing SQL language makes code more complex and harder to maintain. +- Designing a good schema, identifying required indexes etc. takes time +- Some developers are superstitious about sql and relational databases ;) + ==> class attributes * In some classes there are some attributes assigned but never used. (see remove_unused_attributes.patch)